bioRxiv ScienceSearch

Biology subjects

Zeng, P.

Publications and source records attributed to Zeng, P..

7 recordsLinked to original sources

Causal Association between Birth Weight and Adult Diseases: Evidence from a Mendelian Randomisation Analysis

BackgroundIt has long been hypothesized that birth weight has a profound long-term impact on individual predisposition to various diseases at adulthood: a hypothesis commonly referred to as the fetal origins of adult diseases. However, it is not fully clear to what extent the fetal origins of adult diseases hypothesis holds and it is also not completely known what types of adult diseases are causally affected by birth weight. Determining the causal impact of birth weight on various adult diseases through traditional randomised intervention studies is a challenging task.\n\nMethodsMendelian randomisation was employed and multiple genetic variants associated with birth weight were used as instruments to explore the relationship between 21 adult diseases and 38 other complex traits from 37 large-scale genome-wide association studies up to 340,000 individuals of European ancestry. Causal effects of birth weight were estimated using inverse-variance weighted methods. The identified causal relationships between birth weight and adult diseases were further validated through extensive sensitivity analyses and simulations.\n\nResultsAmong the 21 adult diseases, three were identified to be inversely causally affected by birth weight with a statistical significance level passing the Bonferroni corrected significance threshold. The measurement unit of birth weight was defined as its standard deviation (i.e. 488 grams), and one unit lower birth weight was causally related to an increased risk of coronary artery disease (CAD), myocardial infarction (MI), type 2 diabetes (T2D) and BMI-adjusted T2D, with the estimated odds ratios of 1.34 [95% confidence interval (CI) 1.17 - 1.53, p = 1.54E-5], 1.30 (95% CI 1.13 - 1.51, p = 3.31E-4), 1.41 (95% CI 1.15 - 1.73, p = 1.11E-3) and 1.54 (95% CI 1.25 - 1.89, p = 6.07E-5), respectively. All these identified causal associations were robust across various sensitivity analyses that guard against various confounding due to pleiotropy or maternal effects as well as inverse causation. In addition, analysis on 38 additional complex traits found that the inverse causal association between birth weight and CAD/MI/T2D was not likely to be mediated by other risk factors such as blood-pressure related traits and adult weight.\n\nConclusionsThe results suggest that lower birth weight is causally associated with an increased risk of CAD, MI and T2D in later life, supporting the fetal origins of adult diseases hypothesis.

epidemiology

Causal Effects of Blood Lipids on Amyotrophic Lateral Sclerosis: A Mendelian Randomization Study

Amyotrophic lateral sclerosis (ALS) is a late-onset fatal neurodegenerative disorder that is predicted to increase across the globe by ~70% in the following decades. Understanding the disease causal mechanism underlying ALS and identifying modifiable risks factors for ALS hold the key for the development of effective preventative and treatment strategies. Here, we investigate the causal effects of four blood lipid traits that include high density lipoprotein (HDL), low density lipoprotein (LDL), total cholesterol (TC), and triglycerides (TG) on the risk of ALS. By leveraging instrument variables from multiple large-scale genome-wide association studies in both European and East Asian populations, we carry out one of the largest and most comprehensive Mendelian randomization analyses performed to date on the causal relationship between lipids and ALS. Among the four lipids, we found that only LDL is causally associated with ALS and that higher LDL level increases the risk of ALS in both the European and East Asian populations. Specifically, the odds ratio of ALS per one standard deviation (i.e. 39.0 mg/dL) increase of LDL is estimated to be 1.14 (95% CI 1.05 - 1.24, p = 1.38E-3) in the European and population and 1.06 (95% CI 1.00 - 1.12, p = 0.044) in the East Asian population. The identified causal relationship between LDL and ALS is robust with respect to the choice of statistical methods and is validated through extensive sensitivity analyses that guard against various model assumption violations. Our study provides important evidence supporting the causal role of higher LDL on increasing the risk of ALS, paving ways for the development of preventative strategies for reducing the disease burden of ALS across multiple nations.

epidemiology

Jackknife model averaging prediction methods for complex phenotypes with gene expression levels by integrating external pathway information

MotivationIn the past few years many novel prediction approaches have been proposed and widely employed in high dimensional genetic data for disease risk evaluation. However, those approaches typically ignore in model fitting the important group structures or functional classifications that naturally exists in genetic data.\n\nMethodsIn the present study, we applied a novel model averaging approach, called Jackknife Model Averaging Prediction (JMAP), for high dimensional genetic risk prediction while incorporating KEGG pathway information into the model specification. JMAP selects the optimal weights across candidate models by minimizing a cross-validation criterion in a jackknife way. Compared with previous approaches, one of the primary features of JMAP is to allow model weights to vary from 0 to 1 but without the limitation that the summation of weights is equal to one. We evaluated the performance of JMAP using extensive simulation studies and compared it with existing methods. We finally applied JMAP to five real cancer datasets that are publicly available from TCGA.\n\nResultsThe simulations showed that, compared with other existing approaches, JMAP performed best or are among the best methods across a range of scenarios. For example, among 14 out of 16 simulation settings with PVE=0.3, JMAP has an average of 0.075 higher prediction accuracy compared with gsslasso. We further found that in the simulation the model weights for the true candidate models have much smaller chances to be zero compared with those for the null candidate models and are substantially greater in magnitude. In the real data application, JMAP also behaves comparably or better compared with the other methods for both continuous and binary phenotypes. For example, for the COAD, CRC and PAAD data sets, the average gains of predictive accuracy of JMAP are 0.019, 0.064 and 0.052 compared with gsslasso.\n\nConclusionThe proposed method JMAP is a novel method that can provide more accurate phenotypic prediction while incorporating external useful group information.

genomics

Pleiotropic Mapping and Annotation Selection in Genome-wide Association Studies with Pe-nalized Gaussian Mixture Models

MotivationGenome-wide association studies (GWASs) have identified many genetic loci associated with complex traits. A substantial fraction of these identified loci are associated with multiple traits - a phenomena known as pleiotropy. Identification of pleiotropic associations can help characterize the genetic relationship among complex traits and can facilitate our understanding of disease etiology. Effective pleiotropic association mapping requires the development of statistical methods that can jointly model multiple traits with genome-wide SNPs together.\n\nResultsWe develop a joint modeling method, which we refer to as the integrative MApping of Pleiotropic association (iMAP). iMAP models summary statistics from GWASs, uses a multivariate Gaussian distribution to account for phenotypic correlation, simultaneously infers genome-wide SNP association pattern using mixture modeling, and has the potential to reveal causal relationship between traits. Importantly, iMAP integrates a large number of SNP functional annotations to substantially improve association mapping power, and, with a sparsity-inducing penalty, is capable of selecting informative annotations from a large, potentially noninformative set. To enable scalable inference of iMAP to association studies with hundreds of thousands of individuals and millions of SNPs, we develop an efficient expectation maximization algorithm based on an approximate penalized regression algorithm. With simulations and comparisons to existing methods, we illustrate the benefits of iMAP both in terms of high association mapping power and in terms of accurate estimation of genome-wide SNP association patterns. Finally, we apply iMAP to perform a joint analysis of 48 traits from 31 GWAS consortia together with 40 tissue-specific SNP annotations generated from the Roadmap Project. iMAP is freely available at www.xzlab.org/software.html.

bioinformatics

Identifying and exploiting trait-relevant tissues with multiple functional annotations in genome-wide association studies

Genome-wide association studies (GWASs) have identified many disease associated loci, the majority of which have unknown biological functions. Understanding the mechanism underlying trait associations requires identifying trait-relevant tissues and investigating associations in a trait-specific fashion. Here, we extend the widely used linear mixed model to incorporate multiple SNP functional annotations from omics studies with GWAS summary statistics to facilitate the identification of trait-relevant tissues, with which to further construct powerful association tests. Specifically, we rely on a generalized estimating equation based algorithm for parameter inference, a mixture modeling framework for trait-tissue relevance classification, and a weighted sequence kernel association test constructed based on the identified trait-relevant tissues for powerful association analysis. We refer to our analytic procedure as the Scalable Multiple Annotation integration for trait-Relevant Tissue identification and usage (SMART). With extensive simulations, we show how our method can make use of multiple complementary annotations to improve the accuracy for identifying trait-relevant tissues. In addition, our procedure allows us to make use of the inferred trait-relevant tissues, for the first time, to construct more powerful SNP set tests. We apply our method for an in-depth analysis of 43 traits from 28 GWASs using tissue-specific annotations in 105 tissues derived from ENCODE and Roadmap. Our results reveal new trait-tissue relevance, pinpoint important annotations that are informative of trait-tissue relationship, and illustrate how we can use the inferred trait-relevant tissues to construct more powerful association tests in the Wellcome trust case control consortium study.\n\nAuthor SummaryIdentifying trait-relevant tissues is an important step towards understanding disease etiology. Computational methods have been recently developed to integrate SNP functional annotations generated from omics studies to genome-wide association studies (GWASs) to infer trait-relevant tissues. However, two important questions remain to be answered. First, with the increasing number and types of functional annotations nowadays, how do we integrate multiple annotations jointly into GWASs in a trait-specific fashion to take advantage of the complementary information contained in these annotations to optimize the performance of trait-relevant tissue inference? Second, what to do with the inferred trait-relevant tissues? Here, we develop a new statistical method and software to make progress on both fronts. For the first question, we extend the commonly used linear mixed model, with new algorithms and inference strategies, to incorporate multiple annotations in a trait-specific fashion to improve trait-relevant tissue inference accuracy. For the second question, we rely on the close relationship between our proposed method and the widely-used sequence kernel association test, and use the inferred trait-relevant tissues, for the first time, to construct more powerful association tests. We illustrate the benefits of our method through extensive simulations and applications to a wide range of real data sets.

bioinformatics

GIC: A computational method for predicting the essentiality of long noncoding lncRNAs

Measuring the essentiality of genes is critically important in biology and medicine. Some bioinformatic methods have been developed for this issue but none of them can be applied to long noncoding RNAs (lncRNAs), one big class of biological molecules. Here we developed a computational method, GIC (Gene Importance Calculator), which can predict the essentiality of both protein-coding genes and lncRNAs based on RNA sequence information. For identifying the essentiality of protein-coding genes, GIC is competitive with well-established computational scores. More important, GIC showed a high performance for predicting the essentiality of lncRNAs. In an independent mouse lncRNA dataset, GIC achieved an exciting performance (AUC=0.918). In contrast, the traditional computational methods are not applicable to lncRNAs. As a public web server, GIC is freely available at http://www.cuilab.cn/gic/.

bioinformatics

Non-Parametric Genetic Prediction of Complex Traits with Latent Dirichlet Process Regression Models

Using genotype data to perform accurate genetic prediction of complex traits can facilitate genomic selection in animal and plant breeding programs, and can aid in the development of personalized medicine in humans. Because most complex traits have a polygenic architecture, accurate genetic prediction often requires modeling all genetic variants together via polygenic methods. Here, we develop such a polygenic method, which we refer to as the latent Dirichlet process regression model (DPR). DPR is non-parametric in nature, relies on the Dirichlet process to flexibly and adaptively model the effect size distribution, and thus enjoys robust prediction performance across a broad spectrum of genetic architectures. We compare DPR with several commonly used prediction methods with simulations. We further apply DPR to predict gene expressions, to conduct PrediXcan based gene set test, to perform genomic selection of four traits in two species, and to predict eight complex traits in a human cohort.

genetics