bioRxiv ScienceSearch

Biology subjects

Mukherjee, S.

Publications and source records attributed to Mukherjee, S..

At least 19 recordsLinked to original sources

Extrinsic Noise Suppression in Micro RNA Mediated Incoherent Feedforward Loops

MicroRNA mediated incoherent feed forward loops (IFFLs) are recurrent network motifs in mammalian cells and have been a topic of study for their noise rejection and buffering properties. Previous work showed that IFFLs can adapt to varying promoter activity and are less prone to noise than similar circuits without the feed forward loop. Furthermore, it has been shown that microRNAs are better at rejecting extrinsic noise than intrinsic noise. This work studies the biological mechanisms that lead to extrinsic noise rejection for microRNA mediated feed forward network motifs. Specifically, we compare the effects of microRNA-induced mRNA degradation and translational inhibition on extrinsic noise rejection, and identify the parameter regimes where noise is most efficiently rejected. In the case of static extrinsic noise, we find that translational inhibition can expand the regime of extrinsic noise rejection. We then analyze rejection of dynamic extrinsic noise in the case of a single-gene feed forward loop (sgFFL), a special case of the IFFL motif where the microRNA and target mRNA are co-expressed. For this special case, we demonstrate that depending on the time-scale of fluctuations in the extrinsic variable compared to the mRNA and microRNA decay rates, the feed forward loop can both buffer or amplify fluctuations in gene product copy numbers.

systems biology

Evolution of DNA methylation in Papio baboons

Changes in gene regulation have long been thought to play an important role in primate evolution. However, although a number of studies have compared genome-wide gene expression patterns across primate species, fewer have investigated the gene regulatory mechanisms that underlie such patterns, or the relative contribution of drift versus selection. Here, we profiled genome-scale DNA methylation levels from five of the six extant species of the baboon genus Papio (4-14 individuals per species). This radiation presents the opportunity to investigate DNA methylation divergence at both shallow and deeper time scales (380,000 - 1.4 million years). In contrast to studies in human populations, but similar to studies in great apes, DNA methylation profiles clearly mirror genetic and geographic structure. Divergence in DNA methylation proceeds fastest in unannotated regions of the genome and slowest in regions of the genome that are likely more constrained at the sequence level (e.g., gene exons). Both heuristic approaches and Ornstein-Uhlenbeck models suggest that DNA methylation levels at a small set of sites have been affected by positive selection, and that this class is enriched in functionally relevant contexts, including promoters, enhancers, and CpG islands. Our results thus indicate that the rate and distribution of DNA methylation changes across the genome largely mirror genetic structure. However, at some CpG sites, DNA methylation levels themselves may have been a target of positive selection, pointing to loci that could be important in connecting sequence variation to fitness-related traits.

genomics

Diagnostic value of blood gene expression-based classifiers as exemplified for acute myeloid leukemia

ABSTRACTAcute Myeloid Leukemia (AML) is a severe, mostly fatal hematopoietic malignancy. Despite nearly two decades of promising results using gene expression profiling, international recommendations for diagnosis and differential diagnosis of AML remain based on classical approaches including assessment of morphology, immunophenotyping, cytochemistry, and cytogenetics. Concerns about the translation of whole transcriptome profiling include the robustness of derived predictors when taking into account factors such as study- and site-specific effects and whether achievable levels of accuracy are sufficient for practical use. In the present study, we sought to shed light on these issues via a large-scale analysis using machine learning methods applied to a total of 12,029 samples from 105 different studies. Taking advantage of the breadth of data and the now much improved understanding of high-dimensional modeling, we show that AML can be predicted with high accuracy. High-dimensional approaches - in which multivariate signatures are learned directly from genome-wide data with no prior biological knowledge - are highly effective and robust. We explore also the relationship between predictive signatures, differential expression and known AML-related genes. Taken together, our results support the notion that transcriptome assessment could be used as part of an integrated genomic approach in cancer diagnosis and treatment to be implemented early on for diagnosis and differential diagnosis of AML.\n\nOne Sentence SummaryBlood gene expression data and machine learning were used to develop robust and accurate classifiers for diagnosis and differential diagnosis of acute myeloid leukemia based on analysis of more than 12,000 samples derived from more than 100 individual studies

genomics

Genetic data and cognitively-defined late-onset Alzheimer’s disease subgroups

Categorizing people with late-onset Alzheimers disease into biologically coherent subgroups is important for personalized medicine. We evaluated data from five studies (total n=4 050, of whom 2 431 had genome-wide single nucleotide polymorphism (SNP) data). We assigned people to cognitively-defined subgroups on the basis of relative performance in memory, executive functioning, visuospatial functioning, and language at the time of Alzheimers disease diagnosis. We compared genotype frequencies for each subgroup to those from cognitively normal elderly controls. We focused on APOE and on SNPs with p<10-5 and odds ratios more extreme than those previously reported for Alzheimers disease (<0.77 or >1.30). There was substantial variation across studies in the proportions of people in each subgroup. In each study, higher proportions of people with isolated substantial relative memory impairment had [&ge;]1 APOE e4 allele than any other subgroup (overall p= 1.5 x 10-27). Across subgroups, there were 33 novel suggestive loci across the genome with p<10-5 and an extreme OR compared to controls, of which none had statistical evidence of heterogeneity and 30 had ORs in the same direction across all datasets. These data support the biological coherence of cognitively-defined subgroups and nominate novel genetic loci.

genetics

Japanese Encephalitis Virus Infected Microglial Cells Secrete Exosomes Containing let-7a/b that Facilitate Neuronal Damage via Caspase Activation

Extracellular microRNAs (miRNAs) are essential for the cell to cell communication in the healthy and diseased brain. MicroRNAs released from the activated microglia upon neurotropic virus infection may exacerbate CNS damage. Here, we identified let-7a and let-7b (let-7a/b) as the overexpressed miRNAs in Japanese Encephalitis virus (JEV) infected microglia and assessed their role in JEV pathogenesis. We measured the let-7a/b expressions in JEV infected post-mortem human brains, mice brains and in mouse microglial N9 cells by the qRT-PCR and in situ hybridization assay. The interaction between let-7a/b and NOTCH signaling pathway further examined in Toll-like receptor 7 knockdown (TLR7 KD) mice to assess the functions. Exosomes released from JEV infected or let-7a/b mimic transfected N9, and HEK-293 cells were isolated and evaluated their function. We observed an upregulation of let-7a/b in the infected brains as well as in microglia. Knockdown of TLR7 or Inhibition of let-7a/b suppressed the JEV induced NOTCH activation possibly via NF-{kappa}B dependent manner and subsequently, attenuated JEV induced TNF production in microglial cells. Further, exosomes secreted from JEV-infected microglial cells specifically contained let-7a/b. Exosomes overexpressed with let-7a/b were injected into BALB/c mice as well as co-incubated with mouse neuronal (Neuro2a) cells, or primary cortical neuron resulted in caspase activation leading to neuronal damage in the brain. Thus, our results provide evidence for the multifaceted role of let-7a/b miRNAs and unravel the exosomes mediated mechanism for JEV induced pathogenesis.

neuroscience

Label propagation defines signaling networks associated with recurrently mutated cancer genes

Each different tumor type has a distinct profile of genomic perturbations and each of these alterations causes unique changes to cellular homeostasis. Detailed analyses of these changes would reveal downstream effects of genomic alterations, contributing to our understanding of their roles in tumor development and progression. Across a range of tumor types, including bladder, lung, and endometrial carcinoma, we determined genes that are frequently altered in The Cancer Genome Atlas patient populations to study the effects of these alterations on signaling and regulatory pathways. To achieve this, we used a label propagation-based methodology to generate networks from gene expression signatures of mutations. Individual networks offered a comprehensive view of signaling changes represented by gene signatures, which in turn reflect the scope of molecular events that are perturbed in the presence of a given genomic alteration. Comparing different networks to each other revealed commonalities between them and biological pathways distinct genomic alterations converge on, highlighting the critical signaling events tumor dysregulate through multiple mechanisms. Finally, mutations inducing common changes to the signaling network were used to search for genomic markers of drug response, connecting shared perturbations to differential drug response.

bioinformatics

The variations of human miRNAs and Ising like base pairing models

miRNAs are small about 22-base pair long, RNA molecules are of extreme biological importance. Like other longer RNA molecules, messages in miRNAs are encoded by the permutations of only four nucleotide bases represented by A, U, C and G. However, just like words in any language, not all combination of these alphabets make a meaningful word. In fact, we find that the distributions of nucleotides bases in human miRNAs show significant deviation from randomness. First, a miRNA sequence containing four bases are mapped into a binary string with three kinds of classifications according to their chemical properties. Then, we propose a simple nearest neighbor model (Ising model) to understand the statistical variations in human miRNAs.

bioinformatics

Combinations of DIPs and Dprs control organization of olfactory receptor neuron terminals in Drosophila

In Drosophila, 50 classes of olfactory receptor neurons (ORNs) connect to 50 class-specific and uniquely positioned glomeruli in the antennal lobe. Despite the identification of cell surface receptors regulating axon guidance, how ORN axons sort to form 50 stereotypical glomeruli remains unclear. Here we show that the heterophilic cell adhesion proteins, DIPs and Dprs, are expressed in ORNs during glomerular formation. Each ORN class expresses a unique combination of DIPs/dprs, with neurons of the same class expressing interacting partners, suggesting a role in class-specific self-adhesion ORN axons. Analysis of DIP/Dpr expression revealed that ORNS that target neighboring glomeruli have different combinations, and ORNs with very similar DIP/Dpr combinations can project to distant glomeruli in the antennal lobe. Perturbations of DIP/dpr gene function result in local projection defects of ORN axons and glomerular positioning, without altering correct matching of ORNs with their target neurons. Our results suggest that context-dependent differential adhesion through DIP/Dpr combinations regulate self-adhesion and sort ORN axons into uniquely positioned glomeruli.

neuroscience

Dynamic linear models guide design and analysis of microbiota studies within artificial human guts

Artificial gut models provide unique opportunities to study human-associated microbiota. Outstanding questions for these models fundamental biology include the timescales on which microbiota vary and the factors that drive such change. Answering these questions though requires overcoming analytical obstacles like estimating the effects of technical variation on observed microbiota dynamics, as well as the lack of appropriate benchmark datasets. To address these obstacles, we created a modeling framework based on multinomial logistic-normal dynamic linear models (MALLARDs) and performed dense longitudinal sampling of replicate artificial human guts over the course of 1 month. The resulting analyses revealed that when observed on an hourly basis, 76% of community variation could be ascribed to technical noise from sample processing, which could also skew the observed covariation between taxa. Our analyses also supported hypotheses that human gut microbiota fluctuate on sub-daily timescales in the absence of a host and that microbiota can follow replicable trajectories in the presence of environmental driving forces. Finally, multiple aspects of our approach are generalizable and could ultimately be used to facilitate the design and analysis of longitudinal microbiota studies in vivo.

genomics

Unifying mutualism diversity for interpretation and prediction

Coarse-grained rules are widely used in chemistry, physics and engineering. In biology, however, such rules are less common and under-appreciated. This gap can be attributed to the difficulty in establishing general rules to encompass the immense diversity and complexity of biological systems. Even when a rule is established, it is often challenging to map it to mechanistic details and to quantify these details. We here address these challenges on a study of mutualism, an essential type of ecological interaction in nature. Using an appropriate level of abstraction, we deduced a general rule that predicts the outcomes of mutualistic systems, including coexistence and productivity. We further developed a standardized calibration procedure to apply the rule to mutualistic systems without the need to fully elucidate or characterize their mechanistic underpinnings. Our approach consistently provides explanatory and predictive power with various simulated and experimental mutualistic systems. Our strategy can pave the way for establishing and implementing other simple rules for biological systems.

systems biology

A computational framework identifying concordant gene expression-neuropathology associations reveals Complex I as a potential Alzheimer’s disease therapeutic target

Identifying gene expression markers for Alzheimers disease (AD) neuropathology through meta-analysis is a complex undertaking because available data are often from different studies and/or brain regions involving study-specific confounders and/or region-specific biological processes. Here we introduce a novel probabilistic model-based framework, DECODER, leveraging these discrepancies to identify robust biomarkers for complex phenotypes. Our experiments present: (1) DECODERs potential as a general meta-analysis framework widely applicable to various diseases (e.g., AD and cancer) and phenotypes (e.g., Amyloid-{beta} (A{beta}) pathology, tau pathology, and survival), (2) our results from a meta-analysis using 1,746 human brain tissue samples from nine brain regions in three studies -- the largest expression meta-analysis for AD, to our knowledge --, and (3) in vivo validation of identified modifiers of A{beta} toxicity in a transgenic Caenorhabditis elegans model expressing AD-associated A{beta}, which pinpoints mitochondrial Complex I as a critical mediator of proteostasis and a promising pharmacological avenue toward treating AD.

systems biology

Identifying progressive gene network perturbation from single-cell RNA-seq data

Identifying the gene regulatory networks that control development or disease is one of the most important problems in biology. Here, we introduce a computational approach, called PIPER (ProgressIve network PERturbation), to identify the perturbed genes that drive differences in the gene regulatory network across different points in a biological progression. PIPER employs algorithms tailor-made for single cell RNA sequencing (scRNA-seq) data to jointly identify gene networks for multiple progressive conditions. It then performs differential network analysis along the identified gene networks to identify master regulators. We demonstrate that PIPER outperforms state-of-the-art alternative methods on simulated data and is able to predict known key regulators of differentiation on real scRNA-Seq datasets.

bioinformatics

A statistical framework for cross-tissue transcriptome-wide association analysis

Transcriptome-wide association analysis is a powerful approach to studying the genetic architecture of complex traits. A key component of this approach is to build a model to predict (impute) gene expression levels from genotypes from samples with matched genotypes and expression levels in a specific tissue. However, it is challenging to develop robust and accurate imputation models with limited sample sizes for any single tissue. Here, we first introduce a multi-task learning approach to jointly impute gene expression in 44 human tissues. Compared with single-tissue methods, our approach achieved an average 39% improvement in imputation accuracy and generated effective imputation models for an average 120% (range 13%-339%) more genes in each tissue. We then describe a summary statistic-based testing framework that combines multiple single-tissue associations into a single powerful metric to quantify overall gene-trait association at the organism level. When our method, called UTMOST, was applied to analyze genome wide association results for 50 complex traits (Ntotal=4.5 million), we were able to identify considerably more genes in tissues enriched for trait heritability, and cross-tissue analysis significantly outperformed single-tissue strategies (p=1.7e-8). Finally, we performed a cross-tissue genome-wide association study for late-onset Alzheimers disease (LOAD) and replicated our findings in two independent datasets (Ntotal=175,776). In total, we identified 69 significant genes, many of which are novel, leading to novel insights on LOAD etiologies.

genetics

CHC22 Clathrin Diverts GLUT4 from the ER-to-Golgi Intermediate Compartment for Intracellular Sequestration

Glucose Transporter 4 (GLUT4) is sequestered inside muscle and fat, then released by vesicle traffic to the cell surface in response to post-prandial insulin for blood glucose clearance. Here we map the biogenesis of this GLUT4 traffic pathway in humans, which involves clathrin isoform CHC22. We observe that GLUT4 transits through the early secretory pathway more slowly than the constitutively-secreted GLUT1 transporter and localize CHC22 to the endoplasmic-reticulum-to-Golgi-intermediate compartment (ERGIC). CHC22 functions in transport from the ERGIC, as demonstrated by an essential role in forming the replication vacuole of Legionella pneumophila bacteria, which requires ERGIC-derived membrane. CHC22 complexes with ERGIC tether p115, GLUT4 and sortilin and down-regulation of either p115 or CHC22, but not GM130 or sortilin abrogate insulin-responsive GLUT4 release. This indicates CHC22 traffic initiates human GLUT4 sequestration from the ERGIC, and defines a role for CHC22 in addition to retrograde sorting of GLUT4 after endocytic recapture, enhancing pathways for GLUT4 sequestration in humans relative to mice, which lack CHC22.\n\nSummaryBlood glucose clearance relies on insulin-mediated exocytosis of glucose transporter 4 (GLUT4) from sites of intracellular sequestration. We show that in humans, CHC22 clathrin mediates membrane traffic from the ER-to-Golgi Intermediate Compartment, which is needed for GLUT4 sequestration during GLUT4 pathway biogenesis.

cell biology

Phylofactorization - a graph partitioning algorithm to identify phylogenetic scales of ecological data

The problem of pattern and scale is a central challenge in ecology. The problem of scale is central to community ecology, where functional ecological groups are aggregated and treated as a unit underlying an ecological pattern, such as aggregation of \"nitrogen fixing trees\" into a total abundance of a trait underlying ecosystem physiology. With the emergence of massive community ecological datasets, from microbiomes to breeding bird surveys, there is a need to objectively identify the scales of organization pertaining to well-defined patterns in community ecological data.\n\nThe phylogeny is a scaffold for identifying key phylogenetic scales associated with macroscopic patterns. Phylofactorization was developed to objectively identify phylogenetic scales underlying patterns in relative abundance data. However, many ecological data, such as presence-absences and counts, are not relative abundances, yet it is still desireable and informative to identify phylogenetic scales underlying a pattern of interest. Here, we generalize phylofactorization beyond relative abundances to a graph-partitioning algorithm for any community ecological data.\n\nGeneralizing phylofactorization connects many tools from data analysis to phylogenetically-informe analysis of community ecological data. Two-sample tests identify three phylogenetic factors of mammalian body mass which arose during the K-Pg extinction event, consistent with other analyses of mammalian body mass evolution. Projection of data onto coordinates defined by the phylogeny yield a phylogenetic principal components analysis which refines our understanding of the major sources of variation in the human gut microbiome. These same coordinates allow generalized additive modeling of microbes in Central Park soils and confirm that a large clade of Acidobacteria thrive in neutral soils. Generalized linear and additive modeling of exponential family random variables can be performed by phylogenetically-constrained reduced-rank regression or stepwise factor contrasts. We finish with a discussion of how phylofac-torization produces an ecological species concept with a phylogenetic constraint. All of these tools can be implemented with a new R package available online.

ecology

Prior Knowledge And Sampling Model Informed Learning With Single Cell RNA-Seq Data

MotivationSingle cell RNA-seq (scRNA-seq) data contains a wealth of information which has to be inferred computationally from the observed sequencing reads. As the ability to sequence more cells improves rapidly, existing computational tools suffer from three problems. (1) The decreased reads-per-cell implies a highly sparse sample of the true cellular transcriptome. (2) Many tools simply cannot handle the size of the resulting datasets. (3) Prior biological knowledge such as bulk RNA-seq information of certain cell types or qualitative marker information is not taken into account. Here we present UNCURL, a preprocessing framework based on non-negative matrix factorization for scRNA-seq data, that is able to handle varying sampling distributions, scales to very large cell numbers and can incorporate prior knowledge.\n\nResultsWe find that preprocessing using UNCURL consistently improves performance of commonly used scRNA-seq tools for clustering, visualization, and lineage estimation, both in the absence and presence of prior knowledge. Finally we demonstrate that UNCURL is extremely scalable and parallelizable, and runs faster than other methods on a scRNA-seq dataset containing 1.3 million cells.\n\nAvailabilitySource code is available at https://github.com/yjzhang/uncurl_python\n\nContactksreeram@uw.edu, gseelig@uw.edu

bioinformatics

A powerful approach to estimating annotation-stratified genetic covariance using GWAS summary statistics

Despite the success of large-scale genome-wide association studies (GWASs) on complex traits, our understanding of their genetic architecture is far from complete. Jointly modeling multiple traits genetic profiles has provided insights into the shared genetic basis of many complex traits. However, large-scale inference sets a high bar for both statistical power and biological interpretability. Here we introduce a principled framework to estimate annotation-stratified genetic covariance between traits using GWAS summary statistics. Through theoretical and numerical analyses we demonstrate that our method provides accurate covariance estimates, thus enabling researchers to dissect both the shared and distinct genetic architecture across traits to better understand their etiologies. Among 50 complex traits with publicly accessible GWAS summary statistics (Ntotal {approx} 4.5 million), we identified more than 170 pairs with statistically significant genetic covariance. In particular, we found strong genetic covariance between late-onset Alzheimers disease (LOAD) and amyotrophic lateral sclerosis (ALS), two major neurodegenerative diseases, in single-nucleotide polymorphisms (SNPs) with high minor allele frequencies and in SNPs located in the predicted functional genome. Joint analysis of LOAD, ALS, and other traits highlights LOADs correlation with cognitive traits and hints at an autoimmune component for ALS.

genetics

Scaling single cell transcriptomics through split pool barcoding

Constructing an atlas of cell types in complex organisms will require a collective effort to characterize billions of individual cells. Single cell RNA sequencing (scRNA-seq) has emerged as the main tool for characterizing cellular diversity, but current methods use custom microfluidics or microwells to compartmentalize single cells, limiting scalability and widespread adoption. Here we present Split Pool Ligation-based Transcriptome sequencing (SPLiT-seq), a scRNA-seq method that labels the cellular origin of RNA through combinatorial indexing. SPLiT-seq is compatible with fixed cells, scales exponentially, uses only basic laboratory equipment, and costs one cent per cell. We used this approach to analyze 109,069 single cell transcriptomes from an entire postnatal day 5 mouse brain, providing the first global snapshot at this stage of development. We identified 13 main populations comprising different types of neurons, glia, immune cells, endothelia, as well as types in the blood-brain-barrier. Moreover, we resolve substructure within these clusters corresponding to cells at different stages of development. As sequencing capacity increases, SPLiT-seq will enable profiling of billions of cells in a single experiment.

genomics