bioRxiv ScienceSearch

Biology subjects

Weiss, S. T.

Publications and source records attributed to Weiss, S. T..

11 recordsLinked to original sources

Novel Data Transformations for RNA-seq Data Analysis

We propose eight data transformations for RNA-seq data analysis aiming to make the transformed sample mean to be representative of the distribution center since it is not always possible to transform count data to satisfy the normality assumption. Simulation studies showed that limma based on transformed data by using the rv transformation (denoted as limma+rv) performed best compared with limma based on transformed data by using other transformation methods in term of high accuracy and low FNR, while keeping FDR at the nominal level. For large sample size, limma based on transformed data by using the 8 proposed transformation methods had similar performance to limma based on transformed data by using existing transformation methods for equal library size scenarios. Otherwise, limma based on transformed data by using the rv, lv, rv2, or lv2 transformation, or by using the existing voom transformation performed better than limma based on data from other transformation methods. Real data analysis results showed that limma+ l2 performed best, while limma+ rv also had good performance.

bioinformatics

New Statistical Methods for Constructing Robust Differential Correlation Networks

The interplay among microRNAs (miRNAs) plays an important role in the developments of complex human diseases. Co-expression networks can characterize the interactions among miRNAs. Differential correlation network is a powerful tool to investigate the differences of co-expression networks between cases and controls. To construct a differential correlation network, the Fishers Z-transformation test is usually used. However, the Fishers Z-transformation test requires the normality assumption, the violation of which would result in inflated Type I error rate. Several bootstrapping-based improvements for Fishers Z test have been proposed. However, these methods are too computationally intensive to be used to construct differential correlation networks for high-throughput genomic data. In this article, we proposed six novel robust equal-correlation tests that are computationally efficient. The systematic simulation studies and a real microRNA data analysis showed that one of the six proposed tests (ST5) overall performed better than other methods.

bioinformatics

Detecting Differential Variable microRNAs via Model-Based Clustering

Identifying genomic probes (e.g., DNA methylation marks) is becoming a new approach to detect novel genomic risk factors for complex human diseases. The F test is the standard equal-variance test in Statistics. For high-throughput genomic data, the probe-wise F test has been successfully used to detect biologically relevant DNA methylation marks that have different variances between two groups of subjects (e.g., cases vs. controls). In addition to DNA methylation, microRNA is another mechanism of epigenetics. However, to the best of our knowledge, no studies have identified differentially variable (DV) microRNAs. In this article, we proposed a novel model-based clustering to improve the power of the probe-wise F test to detect DV microRNAs. We imposed special structures on covariance matrices for each cluster of microRNAs based on the prior information about the relationship between variance in cases and variance in controls and about the independence among cases and controls. To the best of our knowledge, the proposed method is the first clustering algorithm that aims to detect DV genomic probes. Simulation studies showed that the proposed method outperformed the probe-wise F test and had certain robustness to the violation of the normality assumption. Based on two real datasets about human hepatocellular carcinoma (HCC), we identified 7 DV-only microRNAs (hsa-miR-1826, hsa-miR-191, hsa-miR-194-star, hsa-miR-222, hsa-miR-502-3p, hsa-miR-93, and hsa-miR-99b) using the proposed method, one (hsa-miR-1826) of which has not yet been reported to relate to HCC in the literature.

bioinformatics

Effects of exclusive breastfeeding on infant gut microbiota: a meta-analysis across studies and populations

Literature regarding the differences in gut microbiota between exclusively breastfed (EBF) and non-EBF infants is meager with large variation in methods and results. We performed a meta-analysis of seven studies (a total of 1825 stool samples from 684 infants) to investigate effects of EBF compared to non-EBF on infant gut microbiota across different populations. In the first 6 months of life, overall bacterial diversity, gut microbiota age, relative abundances of Bacteroidetes and Firmicutes and microbial-predicted pathways related to carbohydrate metabolism were consistently increased; while relative abundances of pathways related to lipid, vitamin metabolism and detoxification were decreased in non-EBF vs. EBF infants. The perturbation in microbial-predicted pathways associated with non-EBF was larger in infants delivered by C-section than delivered vaginally. Longer duration of EBF mitigated diarrhea-associated gut microbiota dysbiosis and the effects of EBF persisted after 6 months of age. These consistent findings across vastly different populations suggest that one of the mechanisms of short and long-term benefits of EBF may be alteration in gut microbes.

microbiology

Controllability in an islet specific regulatory network identifies the transcriptional factor NFATC4, which regulates Type 2 Diabetes associated genes

Probing the dynamic control features of biological networks represents a new frontier in capturing the dysregulated pathways in complex diseases. Here, using patient samples obtained from a pancreatic islet transplantation program, we constructed a tissue-specific gene regulatory network and used the control centrality (Cc) concept to identify the high control centrality (HiCc) pathways, which might serve as key pathobiological pathways for Type 2 Diabetes (T2D). We found that HiCc pathway genes were significantly enriched with modest GWAS p-values in the DIAbetes Genetics Replication And Meta-analysis (DIAGRAM) study. We identified variants regulating gene expression (expression quantitative loci, eQTL) of HiCc pathway genes in islet samples. These eQTL genes showed higher levels of differential expression compared to non-eQTL genes in low, medium and high glucose concentrations in rat islets. Among genes with highly significant eQTL evidence, NFATC4 belonged to four HiCc pathways. We asked if the expressions of T2D-associated candidate genes from GWAS and literature are regulated by Nfatc4 in rat islets. Extensive in vitro silencing of Nfatc4 in rat islet cells displayed reduced expression of 16, and increased expression of 4 putative downstream T2D genes. Overall, our approach uncovers the mechanistic connection of NFATC4 with downstream targets including a previously unknown one, TCF7L2, and establishes the HiCc pathways relationship to T2D.

systems biology

Gene co-expression networks in whole blood implicate multiple interrelated molecular pathways in obese asthma

BackgroundAsthmatic children who develop obesity have poorer outcomes compared to those that do not, including poorer control, more severe symptoms, and greater resistance to standard treatment. Gene expression networks are powerful statistical tools for characterizing the underpinnings of human disease that leverage the putative co-regulatory relationships of genes to infer biological pathways altered in disease states.\n\nObjectiveThe aim of this study was to characterize the biology of childhood asthma complicated by adult obesity.\n\nMethodsWe performed weighted gene co-expression network analysis (WGCNA) of gene expression data in whole blood from 514 adult subjects from the Childhood Asthma Management Program (CAMP). We then performed module preservation and association replication analyses in 418 subjects from two independent asthma cohorts (one pediatric and one adult).\n\nResultsWe identified a multivariate model in which four gene co-expression network modules were associated with incident obesity in CAMP (each P < 0.05). The module memberships were enriched for genes in pathways related to platelets, integrins, extracellular matrix, smooth muscle, NF-{kappa}B signaling, and Hedgehog signaling. The network structures of each of the four obese asthma modules were significantly preserved in both replication cohorts (permutation P = 9.999E-05). The corresponding module gene sets were significantly enriched for differential expression in obese subjects in both replication cohorts (each P < 0.05).\n\nConclusionsOur gene co-expression network profiles thus implicate multiple interrelated pathways in the biology of an important endotype of obese asthma.\n\nKey MessagesO_LIWe hypothesized that individuals with asthma complicated by obesity had distinct blood gene expression signatures.\nC_LIO_LIGene co-expression network analysis implicated several inflammatory biological pathways in one form of obese asthma.\nC_LI\n\nCapsule SummaryThis work addresses a knowledge gap about the molecular relationship between asthma and obesity, suggesting that an endotype of obese asthma, known as asthma complicated by obesity, is underpinned by coherent biological mechanisms.\n\nAbbreviations

genomics

On the Stability Landscape of the Human Gut Microbiome: Implications for Microbiome-based Therapies

Understanding how gut microbial species determine their abundances is crucial in developing any microbiome-based therapy. Towards that end, we show that the compositions of our gut microbiota have characteristic and attractive steady states, and hence respond to perturbations in predictable ways. This is achieved by developing a new method to analyze the stability landscape of the human gut microbiome. In order to illustrate the efficacy of our method and its ecological interpretation in terms of asymptotic stability, this novel method is applied to various human cohorts, including large cross-sectional studies, long longitudinal studies with frequent sampling, and perturbation studies via fecal microbiota transplantation, antibiotic and probiotic treatments. These findings will facilitate future ecological modeling efforts in human microbiome research. Moreover, the method allows for the prediction of the compositional shift of the gut microbiome during the fecal microbiota transplantation process. This result holds promise for translational applications, such as, personalized donor selection when performing fecal microbiota transplantations.\n\nOne Sentence SummaryA new method for analyzing the stability landscape of the human gut microbiome and predicting its steady-state composition is developed.

ecology

Deciphering Functional Redundancy in the Human Microbiome

Although the taxonomic composition of the human microbiome varies tremendously across individuals, its gene composition or functional capacity is highly conserved1-5---implying an ecological property known as functional redundancy. Such functional redundancy is thought to underlie the stability and resilience of the human microbiome6,7, but its origin is elusive. Here, we investigate the basis for functional redundancy in the human microbiome by analyzing its genomic content network --- a bipartite graph that links microbes to the genes in their genomes. We show that this network exhibits several topological features, such as highly nested structure and fat-tailed gene degree distribution, which favor high functional redundancy. To explain the origins of these topological features, we develop a simple genome evolution model that explicitly considers selection pressure, and the processes of gene gain and loss, and horizontal gene transfer. We find that moderate selection pressure and high horizontal gene transfer rate are necessary to generate genomic content networks with both highly nested structure and fat-tailed gene degree distribution, and consequently favor high functional redundancy. These findings provide insights into the relationships between structure and function in complex microbial communities. This work elucidates the potential ecological and evolutionary processes that create and maintain functional redundancy in the human microbiome and contribute to its resilience.

microbiology

Mapping the ecological networks of microbial communities from steady-state data

Microbes form complex and dynamic ecosystems that play key roles in the health of the animals and plants with which they are associated. Such ecosystems are often represented by a directed, signed and weighted ecological network, where nodes represent microbial taxa and edges represent ecological interactions. Inferring the underlying ecological networks of microbial communities is a necessary step towards understanding their assembly rules and predicting their dynamical response to external stimuli. However, current methods for inferring such networks require assuming a particular population dynamics model, which is typically not known a priori. Moreover, those methods require fitting longitudinal abundance data, which is not readily available, and often does not contain the variation that is necessary for reliable inference. To overcome these limitations, here we develop a new method to map the ecological networks of microbial communities using steady-state data. Our method can qualitatively infer the inter-taxa interaction types or signs (positive, negative or neutral) without assuming any particular population dynamics model. Additionally, when the population dynamics is assumed to follow the classic Generalized Lotka-Volterra model, our method can quantitatively infer the inter-taxa interaction strengths and intrinsic growth rates. We systematically validate our method using simulated data, and then apply it to four experimental datasets of microbial communities. Our method offers a novel framework to infer microbial interactions and reconstruct ecological networks, and represents a key step towards reliable modeling of complex, real-world microbial communities, such as the human gut microbiota.

ecology

A Novel Nasal Brush-based Classifier of Asthma Identified by Machine Learning Analysis of Nasal RNA Sequence Data

Asthma is a common, under-diagnosed disease affecting all ages. We sought to identify a nasal brush-based classifier of mild/moderate asthma. 190 subjects with mild/moderate asthma and controls underwent nasal brushing and RNA sequencing of nasal samples. A machine learning-based pipeline identified an asthma classifier consisting of 90 genes interpreted via an L2-regularized logistic regression classification model. This classifier performed with strong predictive value and sensitivity across eight test sets, including (1) a test set of independent asthmatic and control subjects profiled by RNA sequencing (positive and negative predictive values of 1.00 and 0.96, respectively; AUC of 0.994), (2) two independent case-control cohorts of asthma profiled by microarray, and (3) five cohorts with other respiratory conditions (allergic rhinitis, upper respiratory infection, cystic fibrosis, smoking), where the classifier had a low to zero misclassification rate. Following validation in large, prospective cohorts, this classifier could be developed into a nasal biomarker of asthma.

systems biology

Whole Genome Sequencing of Pharmacogenetic Drug Response in Racially and Ethnically Diverse Children with Asthma

Asthma is the most common chronic disease of children, with significant racial/ethnic differences in prevalence, morbidity, mortality and therapeutic response. Albuterol, a bronchodilator medication, is the first-line therapy for asthma treatment worldwide. We performed the largest whole genome sequencing (WGS) pharmacogenetics study to date using data from 1,441 minority children with asthma who had extremely high or low bronchodilator drug response (BDR). We identified population-specific and shared pharmacogenetic variants associated with BDR, including genome-wide significant (p < 3.53 x 10-7) and suggestive (p < 7.06 x 10-6) loci near genes previously associated with lung capacity (DNAH5), immunity (NFKB1 and PLCB1), and {beta}-adrenergic signaling pathways (ADAMTS3 and COX18). Functional analyses centered on NFKB1 revealed potential regulatory function of our BDR-associated SNPs in bronchial smooth muscle cells. Specifically, these variants are in linkage disequilibrium with SNPs in a functionally active enhancer, and are also expression quantitative trait loci (eQTL) for a neighboring gene, SLC39A8. Given the lack of other asthma study populations with WGS data on minority children, replication of our rare variant associations is infeasible. We attempted to replicate our common variant findings in five independent studies with GWAS data. The age-specific associations previously found in asthma and asthma-related traits suggest that the over-representation of adults in our replication populations may have contributed to our lack of statistical replication, despite the functional relevance of the NFKB1 variants demonstrated by our functional assays. Our study expands the understanding of pharmacogenetic analyses in racially/ethnically diverse populations and advances the foundation for precision medicine in at-risk and understudied minority populations.\n\nAUTHOR SUMMARYAsthma is the most common chronic disease among children. Albuterol, a bronchodilator medication, is the first-line therapy for asthma treatment worldwide. In the U.S., asthma prevalence is the highest among Puerto Ricans, intermediate among African Americans and lowest in Whites and Mexicans. Asthma disparities extend to mortality, which is four- to five-fold higher in Puerto Ricans and African Americans compared to Mexicans [1]. Puerto Ricans and African Americans, the populations with the highest asthma prevalence and death rate, also have the lowest albuterol bronchodilator drug response (BDR). We conducted the largest pharmacogenetic study using whole genome sequencing data from 1,441 minority children with asthma who had extremely high or low albuterol bronchodilator drug response. We identified population-specific and shared pharmacogenetic variants associated with BDR. Our findings help inform the direction of future development of asthma medications and our study advances the foundation of precision medicine for at-risk, yet understudied, racially/ethnically diverse populations.

genetics