bioRxiv Science⌕ Search

Biology subjects

Papanicolaou, G.

Publications and source records attributed to Papanicolaou, G..

3 recordsLinked to original sources

Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations

Transcriptome prediction models built with data from European-descent individuals are less accurate when applied to different populations because of differences in linkage disequilibrium patterns and allele frequencies. We hypothesized methods that leverage shared regulatory effects across different conditions, in this case, across different populations may improve cross-population transcriptome prediction. To test this hypothesis, we made transcriptome prediction models for use in transcriptome-wide association studies (TWAS) using different methods (Elastic Net, Joint-Tissue Imputation (JTI), Matrix eQTL, Multivariate Adaptive Shrinkage in R (MASHR), and Transcriptome-Integrated Genetic Association Resource (TIGAR)) and tested their out-of-sample transcriptome prediction accuracy in population-matched and cross-population scenarios. Additionally, to evaluate model applicability in TWAS, we integrated publicly available multi-ethnic genome-wide association study (GWAS) summary statistics from the Population Architecture using Genomics and Epidemiology Study (PAGE) and Pan-UK Biobank with our developed transcriptome prediction models. In regard to transcriptome prediction accuracy, MASHR models performed better or the same as other methods in both population-matched and cross-population transcriptome predictions. Furthermore, in multi-ethnic TWAS, MASHR models yielded more discoveries that replicate in both PAGE and PanUKBB across all methods analyzed, including loci previously mapped in GWAS and new loci previously not found in GWAS. Overall, our study demonstrates the importance of using methods that benefit from different populations effect size estimates in order to improve TWAS for multi-ethnic or underrepresented populations.

genomics↗

The functional impact of rare variation across the regulatory cascade

Each human genome has tens of thousands of rare genetic variants; however, identifying impactful rare variants remains a major challenge. We demonstrate how use of personal multi-omics can enable identification of impactful rare variants by using the Multi-Ethnic Study of Atherosclerosis (MESA) which included several hundred individuals with whole genome sequencing, transcriptomes, methylomes, and proteomes collected across two time points, ten years apart. We evaluated each multi-omic phenotypes ability to separately and jointly inform functional rare variation. By combining expression and protein data, we observed rare stop variants 62x and rare frameshift variants 216x as frequently as controls, compared to 13x to 27x for expression or protein effects alone. We developed a Bayesian hierarchical model to prioritize specific rare variants underlying multi-omic signals across the regulatory cascade. With this approach, we identified rare variants that exhibited large effect sizes on multiple complex traits including height, schizophrenia, and Alzheimers disease.

genomics↗

Protein prediction for trait mapping in diverse populations

Genetically regulated gene expression has helped elucidate the biological mechanisms underlying complex traits. Improved high-throughput technology allows similar interrogation of the genetically regulated proteome for understanding complex trait mechanisms. Here, we used the Trans-omics for Precision Medicine (TOPMed) Multi-omics pilot study, which comprises data from Multi-Ethnic Study of Atherosclerosis (MESA), to optimize genetic predictors of the plasma proteome for genetically regulated proteome-wide association studies (PWAS) in diverse populations. We built predictive models for protein abundances using data collected in TOPMed MESA, for which we have measured 1,305 proteins by a SOMAscan assay. We compared predictive models built via elastic net regression to models integrating posterior inclusion probabilities estimated by fine-mapping SNPs prior to elastic net. In order to investigate the transferability of predictive models across ancestries, we built protein prediction models in all four of the TOPMed MESA populations, African American (n=183), Chinese (n=71), European (n=416), and Hispanic/Latino (n=301), as well as in all populations combined. As expected, fine-mapping produced more significant protein prediction models, especially in African ancestries populations, potentially increasing opportunity for discovery. When we tested our TOPMed MESA models in the independent European INTERVAL study, fine-mapping improved cross-ancestries prediction for some proteins. Using GWAS summary statistics from the Population Architecture using Genomics and Epidemiology (PAGE) study, which comprises ~50,000 Hispanic/Latinos, African Americans, Asians, Native Hawaiians, and Native Americans, we applied S-PrediXcan to perform PWAS for 28 complex traits. The most protein-trait associations were discovered, colocalized, and replicated in large independent GWAS using proteome prediction model training populations with similar ancestries to PAGE. At current training population sample sizes, performance between baseline and fine-mapped protein prediction models in PWAS was similar, highlighting the utility of elastic net. Our predictive models in diverse populations are publicly available for use in proteome mapping methods at https://doi.org/10.5281/zenodo.4837328. Author summaryGene regulation is a critical mechanism underlying complex traits. Transcriptome-wide association studies (TWAS) have helped elucidate potential mechanisms because each association connects a gene rather than a variant to the complex trait. Like genome-wide association studies (GWAS), most TWAS are still conducted exclusively in populations of European ancestry, which misses the opportunity to test the full spectrum of human genetic variation for associations with complex traits. Here, move beyond the transcriptome and because protein measurement assays are growing to allow interrogation of the proteome, we use data from TOPMed MESA to develop genetic predictors of protein abundance in diverse ancestry populations. We compare model-building strategies with the goal of providing the best resource for protein association discovery with available data. We demonstrate how these prediction models can be used to perform proteome-wide association studies (PWAS) in diverse populations. We show the most protein-trait associations were discovered, colocalized, and replicated in independent cohorts using proteome prediction model training populations with similar ancestries to individuals in the GWAS. We shared our protein prediction models and performance statistics publicly to facilitate future proteome mapping studies in diverse populations.

genomics↗