bioRxiv ScienceSearch

Biology subjects

Abecasis, G.

Publications and source records attributed to Abecasis, G..

11 recordsLinked to original sources

Exploring Various Polygenic Risk Scores for Basal Cell Carcinoma, Cutaneous Squamous Cell Carcinoma and Melanoma in the Phenomes of the Michigan Genomics Initiative and the UK Biobank

Polygenic risk scores (PRS) are designed to serve as a single summary measure, condensing information from a large number of genetic variants associated with a disease. They have been used for stratification and prediction of disease risk. The construction of a PRS often depends on the purpose of the study, the available data/summary estimates, and the underlying genetic architecture of a disease. In this paper, we consider several choices for constructing a PRS using summary data obtained from various publicly-available sources including the UK Biobank and evaluate their abilities to predict outcomes derived from electronic health records (EHR). Weexamine the three most common skin cancer subtypes in the USA: basal cellcarcinoma, cutaneous squamous cell carcinoma, and melanoma. The genetic risk profiles of subtypes may consist of both shared and unique elements and we construct PRS to understand the common versus distinct etiology. This study is conducted using data from 30,702 unrelated, genotyped patients of recent European descent from the Michigan Genomics Initiative (MGI), a longitudinal biorepository effort within Michigan Medicine. Using these PRS for various skin cancer subtypes, we conduct a phenome-wide association study (PheWAS) within the MGI data to evaluate their association with secondary traits. PheWAS results are then replicated using population-based UK Biobank data. We develop an accompanying visual catalog called PRSweb that provides detailed PheWAS results and allows users to directly compare different PRS construction methods. The results of this study can provide guidance regarding PRS construction in future PRS-PheWAS studies using EHR data involving disease subtypes.\n\nAuthor summaryIn the study of genetically complex diseases, polygenic risk scores synthesize information from multiple genetic risk factors to provide insight into a patients risk of developing a disease based on his/her genetic profile. These risk scores can be explored in conjunction with health and disease information available in the electronic medical records. They may be associated with diseases that may be related to or precursors of the underlying disease of interest. Limited work is available guiding risk score construction when the goal is to identify associations across the medical phenome. In this paper, we compare different polygenic risk score construction methods in terms of their relationships with the medical phenome. We further propose methods for using these risk scores to decouple the shared and unique genetic profiles of related diseases and to explore related diseases shared and unique secondary associations. Leveraging and harnessing the rich data resources of the Michigan Genomics Initiative, a biorepository effort at Michigan Medicine, and the larger population-based UK Biobank study, we investigated the performance of genetic risk profiling methods for the three most common types of skin cancer: melanoma, basal cell carcinoma and squamous cell carcinoma.

genetics

PROTEIN-CODING VARIANTS IMPLICATE NOVEL GENES RELATED TO LIPID HOMEOSTASIS CONTRIBUTING TO BODY FAT DISTRIBUTION

Body fat distribution is a heritable risk factor for a range of adverse health consequences, including hyperlipidemia and type 2 diabetes. To identify protein-coding variants associated with body fat distribution, assessed by waist-to-hip ratio adjusted for body mass index, we analyzed 228,985 predicted coding and splice site variants available on exome arrays in up to 344,369 individuals from five major ancestries for discovery and 132,177 independent European-ancestry individuals for validation. We identified 15 common (minor allele frequency, MAF[&ge;]5%) and 9 low frequency or rare (MAF<5%) coding variants that have not been reported previously. Pathway/gene set enrichment analyses of all associated variants highlight lipid particle, adiponectin level, abnormal white adipose tissue physiology, and bone development and morphology as processes affecting fat distribution and body shape. Furthermore, the cross-trait associations and the analyses of variant and gene function highlight a strong connection to lipids, cardiovascular traits, and type 2 diabetes. In functional follow-up analyses, specifically in Drosophila RNAi-knockdown crosses, we observed a significant increase in the total body triglyceride levels for two genes (DNAH10 and PLXND1). By examining variants often poorly tagged or entirely missed by genome-wide association studies, we implicate novel genes in fat distribution, stressing the importance of interrogating low-frequency and protein-coding variants.

genetics

emeraLD: Rapid Linkage Disequilibrium Estimation with Massive Data Sets

SummaryEstimating linkage disequilibrium (LD) is essential for a wide range of summary statistics-based association methods for genome-wide association studies (GWAS). Large genetic data sets, e.g. the TOPMed WGS project and UK Biobank, enable more accurate and comprehensive LD estimates, but increase the computational burden of LD estimation. Here, we describe emeraLD (Efficient Methods for Estimation and Random Access of LD), a computational tool that leverages sparsity and haplotype structure to estimate LD orders of magnitude faster than existing tools.\n\nAvailability and ImplementationemeraLD is implemented in C++, and is open source under GPLv3. Source code, documentation, an R interface, and utilities for analysis of summary statistics are freely available at http://github.com/statgen/emeraLD\n\nContactcorbinq@umich.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Genome-wide analysis yields new loci associating with aortic valve stenosis

Aortic valve stenosis (AS) is the most common valvular heart disease, characterized by a thickened and calcified valve causing left ventricular outflow obstruction. Severe AS is a significant cause of morbidity and mortality, affecting approximately 5% of those over 70 years of age1,2,3. Little is known about the genetics of AS, although recently a variant at the LPA locus4 and a rare MYH6 missense variant were found to associate with AS5. We report a large genome-wide association study (GWAS) with a follow-up in up to 7,307 AS cases and 801,073 controls. We identified two new AS loci, on chromosome 1p21 near PALMD (rs7543130; OR=1.20, P=1.2x10-22) and on chromosome 2q22 in TEX41 (rs1830321; OR=1.15, P=1.8x10-13). Rs7543130 also associates with bicuspid aortic valve (BAV) (OR=1.28, P=6.6x10-10) and aortic root diameter (P=1.30x10-8) and rs1830321 associates with BAV (OR=1.12, P=5.3x10-3 and coronary artery disease (CAD) (OR=1.05, P=9.3x10-5). These results indicate that AS is partly rooted in the same processes as cardiac development and atherosclerosis.

genetics

Improved Score Statistics for Meta-analysis in Single-variant and Gene-level Association Studies

Meta-analysis is now an essential tool for genetic association studies, allowing these to combine large studies and greatly accelerating the pace of genetic discovery. Although the standard meta-analysis methods perform equivalently as the more cumbersome joint analysis under ideal settings, they result in substantial power loss under unbalanced settings with various case-control ratios. Here, we investigate why the standard meta-analysis methods lose power under unbalanced settings, and further propose a novel meta-analysis method that performs as efficiently as joint analysis under general settings. Our proposed method can accurately approximate the score statistics obtainable by joint analysis, for both linear and logistic regression models, with and without covariates. In addition, we propose a novel approach to adjust for population stratification by correcting for known population structures through minor allele frequencies (MAFs). In the simulated gene-level association studies under unbalanced settings, our method recovered up to 85% power loss caused by the standard method. We further showed the power gain of our method in gene-level association studies with 26 unbalanced real studies of Age-related Macular Degeneration (AMD). In addition, we took the meta-analysis of three studies of type 2 diabetes (T2D) as an example to discuss the challenges of meta-analyzing multi-ethnic samples. In summary, we propose improved single-variant score statistics in meta-analysis, requiring \"accurate\" population-specific MAFs for multi-ethnic studies. These improved score statistics can be used to construct both single-variant and gene-level association studies, providing a useful framework for ensuring well-powered, convenient, cross-study analyses.

genetics

Narrow-sense heritability estimation of complex traits using identity-by-descent information.

Heritability is a fundamental parameter in genetics. Traditional estimates based on family or twin studies can be biased due to shared environmental or non-additive genetic variance. Alternatively, those based on genotyped or imputed variants typically underestimate narrow-sense heritability contributed by rare or otherwise poorly-tagged causal variants. Identical-by-descent (IBD) segments of the genome share all variants between pairs of chromosomes except new mutations that have arisen since the last common ancestor. Therefore, relating phenotypic similarity to degree of IBD sharing among classically unrelated individuals is an appealing approach to estimating the near full additive genetic variance while avoiding biases that can occur when modeling close relatives. We applied an IBD-based approach (GREML-IBD) to estimate heritability in unrelated individuals using phenotypic simulation with thousands of whole genome sequences across a range of stratification, polygenicity levels, and the minor allele frequencies of causal variants (CVs). IBD-based heritability estimates were unbiased when using unrelated individuals, even for traits with extremely rare CVs, but stratification led to strong biases in IBD-based heritability estimates with poor precision. We used data on two traits in ~120,000 people from the UK Biobank to demonstrate that, depending on the trait and possible confounding environmental effects, GREML-IBD can be applied successfully to very large genetic datasets to infer the contribution of very rare variants lost using other methods. However, we observed apparent biases in this real data that were not predicted from our simulation, suggesting that more work may be required to understand factors that influence IBD-based estimates.

genetics

Whole Genome Sequencing in Psychiatric Disorders: the WGSPD Consortium

As technology advances, whole genome sequencing (WGS) is likely to supersede other genotyping technologies. The rate of this change depends on its relative cost and utility. Variants identified uniquely through WGS may reveal novel biological pathways underlying complex disorders and provide high-resolution insight into when, where, and in which cell type these pathways are affected. Alternatively, cheaper and less computationally intensive approaches may yield equivalent insights. Understanding the role of rare variants in the noncoding gene-regulating genome, through pilot WGS projects, will be critical to determine which of these two extremes best represents reality. With large cohorts, well-defined risk loci, and a compelling need to understand the underlying biology, psychiatric disorders have a role to play in this preliminary WGS assessment. The WGSPD consortium will integrate data for 18,000 individuals with psychiatric disorders, beginning with autism spectrum disorder, schizophrenia, bipolar disorder, and major depressive disorder, along with over 150,000 controls.

genomics

Comparison of methods that use whole genome data to estimate the heritability and genetic architecture of complex traits.

Heritability, h2, is a foundational concept in genetics, critical to understanding the genetic basis of complex traits. Recently-developed methods that estimate heritability from genotyped SNPs, h2 SNP, explain substantially more genetic variance than genome-wide significant loci, but less than classical estimates from twins and families. However, h2SNP estimates have yet to be comprehensively compared under a range of genetic architectures, making it difficult to draw conclusions from sometimes conflicting published estimates. Here, we used thousands of real whole genome sequences to simulate realistic phenotypes under a variety of genetic architectures, including those from very rare causal variants. We compared the performance of ten methods across different types of genotypic data (commercial SNP array positions, whole genome sequence variants, and imputed variants) and under differing causal variant frequencies, levels of stratification, and relatedness thresholds. These results provide guidance in interpreting past results and choosing optimal approaches for future studies. We then chose two methods (GREML-MS and GREML-LDMS) that best estimated overall h2SNP and the causal variant frequency spectra to six phenotypes in the UK Biobank using imputed genome-wide variants. Our results suggest that as imputation reference panels become larger and more diverse, estimates of the frequency distribution of causal variants will become increasingly unbiased and the vast majority of trait narrow-sense heritability will be accounted for.

genetics

Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome

A primary goal of The Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to develop an African Diaspora Power Chip (ADPC), a genotyping array consisting of tagging SNPs, useful in comprehensively identifying African specific genetic variation. This array is designed based on the novel variation identified in 642 CAAPA samples of African ancestry with high coverage whole genome sequence data (~30x depth). This novel variation extends the pattern of variation catalogued in the 1000 Genomes and Exome Sequencing Projects to a spectrum of populations representing the wide range of West African genomic diversity. These individuals from CAAPA also comprise a large swath of the African Diaspora population and incorporate historical genetic diversity covering nearly the entire Atlantic coast of the Americas. Here we show the results of designing and producing such a microchip array. This novel array covers African specific variation far better than other commercially available arrays, and will enable better GWAS analyses for researchers with individuals of African descent in their study populations. A recent study1 cataloging variation in continental African populations suggests this type of African-specific genotyping array is both necessary and valuable for facilitating large-scale GWAS in populations of African ancestry.

genomics

Imputation aware tag SNP selection to improve power for multi-ethnic association studies

The emergence of very large cohorts in genomic research has facilitated a focus on genotype-imputation strategies to power rare variant association. Consequently, a new generation of genotyping arrays are being developed designed with tag single nucleotide polymorphisms (SNPs) to improve rare variant imputation. Selection of these tag SNPs poses several challenges as rare variants tend to be continentally-or even population-specific and reflect fine-scale linkage disequilibrium (LD) structure impacted by recent demographic events. To explore the landscape of tag-able variation and guide design considerations for large-cohort and biobank arrays, we developed a novel pipeline to select tag SNPs using the 26 population reference panel from Phase of the 1000 Genomes Project. We evaluate our approach using leave-one-out internal validation via standard imputation methods that allows the direct comparison of tag SNP performance by estimating the correlation of the imputed and real genotypes for each iteration of potential array sites. We show how this approach allows for an assessment of array design and performance that can take advantage of the development of deeper and more diverse sequenced reference panels. We quantify the impact of demography on tag SNP performance across populations and provide population-specific guidelines for tag SNP selection. We also examine array design strategies that target single populations versus multi-ethnic cohorts, and demonstrate a boost in performance for the latter can be obtained by prioritizing tag SNPs that contribute information across multiple populations simultaneously. Finally, we demonstrate the utility of improved array design to provide meaningful improvements in power, particularly in trans-ethnic studies. The unified framework presented will enable investigators to make informed decisions for the design of new arrays, and help empower the next phase of rare variant association for global health.

genomics

A scalable Bayesian method for integrating functional information in genome-wide association studies

Although genome-wide association studies (GWASs) have identified many risk loci for complex traits and common diseases, most of the identified associations reside in noncoding regions and have unknown biological functions. Recent genomic sequencing studies have produced a rich resource of annotations that help characterize the function of genetic variants. Integrative analysis that incorporates these functional annotations into GWAS can help elucidate the biological mechanisms underlying the identified associations and help prioritize causal-variants. Here, we develop a novel, flexible Bayesian variable selection model with efficient computational techniques for such integrative analysis. Different from previous approaches, our method models the effect-size distribution and probability of causality for variants with different annotations and jointly models genome-wide variants to account for linkage disequilibrium (LD), thus prioritizing associations based on the quantification of the annotations and allowing for multiple causal-variants per locus. Our efficient computational algorithm dramatically improves both computational speed and posterior sampling convergence by taking advantage of the block-wise LD structures of human genomes. With simulations, we show that our method accurately quantifies the functional enrichment and performs more powerful for identifying true causal-variants than several competing methods. The power gain brought up by our method is especially apparent in cases when multiple causal-variants in LD reside in the same locus. We also apply our method for an in-depth GWAS of age-related macular degeneration with 33,976 individuals and 9,857,286 variants. We find the strongest enrichment for causality among non-synonymous variants (54x more likely to be causal, 1.4x larger effect-sizes) and variants in active promoter (7.8x more likely, 1.4x larger effect-sizes), as well as identify 5 potentially novel loci in addition to the 32 known AMD risk loci. In conclusion, our method is shown to efficiently integrate functional information in GWASs, helping identify causal variants and underlying biology.\n\nAuthor summaryWe propose a novel Bayesian hierarchical model to account for linkage disequilibrium (LD) and multiple functional annotations in GWAS, paired with an expectation-maximization Markov chain Monte Carlo (EM-MCMC) computational algorithm to jointly analyze genome-wide variants. Our method improves the MCMC convergence property to ensure accurate Bayesian inference of the quantifications of the functional enrichment pattern and fine-mapped association results. By applying our method to the real GWAS of age-related macular degeneration (AMD) with various functional annotations (i.e., gene-based, regulatory, and chromatin states), we find that the variants of non-synonymous, coding, and active promoter annotations have the highest causal probability and the largest effect-sizes. In addition, our method produces fine-mapped association results in the identified risk loci, two of which are shown as examples (C2/CFB/SKIV2L and C3) with justifications by haplotype analysis, model comparison, and conditional analysis. Therefore, we believe our integrative method will be useful for quantifying the enrichment pattern of functional annotations in GWAS, and then prioritizing associations with respect to the learned functional enrichment pattern.

genetics