bioRxiv ScienceSearch

Biology subjects

Beaty, T. H.

Publications and source records attributed to Beaty, T. H..

7 recordsLinked to original sources

Integration of Molecular Interactome and Targeted Interaction Analysis to Identify a COPD Disease Network Module

The polygenic nature of complex diseases offers potential opportunities to utilize network-based approaches that leverage the comprehensive set of protein-protein interactions (the human interactome) to identify new genes of interest and relevant biological pathways. However, the incompleteness of the current human interactome prevents it from reaching its full potential to extract network-based knowledge from gene discovery efforts, such as genome-wide association studies, for complex diseases like chronic obstructive pulmonary disease (COPD). Here, we provide a framework that integrates the existing human interactome information with new experimental protein-protein interaction data for FAM13A, one of the most highly associated genetic loci to COPD, to find a more comprehensive disease network module. We identified an initial disease network neighborhood by applying a random-walk method. Next, we developed a network-based closeness approach (CAB) that revealed 9 out of 96 FAM13A interacting partners identified by affinity purification assays were significantly close to the initial network neighborhood. Moreover, compared to a similar method (local radiality), the CAB approach predicts low-degree genes as potential candidates. The candidates identified by the network-based closeness approach were combined with the initial network neighborhood to build a comprehensive disease network module (163 genes) that was enriched with genes differentially expressed between controls and COPD subjects in alveolar macrophages, lung tissue, sputum, blood, and bronchial brushing datasets. Overall, we demonstrate an approach to find disease-related network components using new laboratory data to overcome incompleteness of the current interactome.

systems biology

Expanded genetic landscape of chronic obstructive pulmonary disease reveals heterogeneous cell type and phenotype associations

Chronic obstructive pulmonary disease (COPD) is the leading cause of respiratory mortality worldwide. Genetic risk loci provide novel insights into disease pathogenesis. To broaden COPD genetic risk loci discovery and identify cell type and phenotype associations we performed a genome-wide association study in 35,735 cases and 222,076 controls from the UK Biobank and additional studies from the International COPD Genetics Consortium. We identified 82 loci with P value < 5x10-8; 47 were previously described in association with either COPD or population-based lung function. Of the remaining 35 novel loci, 13 were associated with lung function in 79,055 individuals from the SpiroMeta consortium. Using gene expression and regulation data, we identified enrichment for loci in lung tissue, smooth muscle and alveolar type II cells. We found 9 shared genomic regions between COPD and asthma and 5 between COPD and pulmonary fibrosis. COPD genetic risk loci clustered into groups of quantitative imaging features and comorbidity associations. Our analyses provide further support to the genetic susceptibility and heterogeneity of COPD.

genetics

New genetic signals for lung function highlight pathways and pleiotropy, and chronic obstructive pulmonary disease associations across multiple ancestries

Reduced lung function predicts mortality and is key to the diagnosis of COPD. In a genome-wide association study in 400,102 individuals of European ancestry, we define 279 lung function signals, one-half of which are new. In combination these variants strongly predict COPD in deeply-phenotyped patient populations. Furthermore, the combined effect of these variants showed generalisability across smokers and never-smokers, and across ancestral groups. We highlight biological pathways, known and potential drug targets for COPD and, in phenome-wide association studies, autoimmune-related and other pleiotropic effects of lung function associated variants. This new genetic evidence has potential to improve future preventive and therapeutic strategies for COPD.

genomics

Inferring Disease Risk Genes from Sequencing Data in Multiplex Pedigrees Through Sharing of Rare Variants

We previously demonstrated how sharing of rare variants (RVs) in distant affected relatives can be used to identify variants causing a complex and heterogeneous disease. This approach tested whether single RVs were shared by all sequenced affected family members. However, as with other study designs, joint analysis of several RVs (e.g. within genes) is sometimes required to obtain sufficient statistical power. Further, phenocopies can lead to false negatives for some causal RVs if complete sharing among affecteds is required. Here we extend our methodology (Rare Variant Sharing, RVS) to address these issues. Specifically, we introduce gene-based analyses, refine RV definition based on haplotypes, and introduce a partial sharing test based on RV sharing probabilities for subsets of affected family members. RVS also has the desirable features of not requiring external estimates of variant frequency or control samples, provides functionality to assess and address violations of key assumptions, and is available as open source software for genome-wide analysis. Simulations including phenocopies, based on the families of an oral cleft study, revealed the partial and complete sharing versions of RVS achieved similar statistical power compared to alternative methods (RareIBD and the Gene-Based Segregation Test), and had superior power compared to the pedigree Variant Annotation, Analysis and Search Tool (pVAAST) linkage statistic. In studies of multiplex cleft families, analysis of rare single nucleotide variants in the exome of 151 affected relatives from 54 families revealed no significant excess sharing in any one gene, but highlighted different patterns of sharing revealed by the complete and partial sharing tests.

genetics

Detection of de novo copy number deletions from targeted sequencing of trios

De novo copy number deletions have been implicated in many diseases, but there is no formal method to date however that identifies de novo deletions in parent-offspring trios from capture-based sequencing platforms. We developed Minimum Distance for Targeted Sequencing (MDTS) to fill this void. MDTS has similar sensitivity (recall), but a much lower false positive rate compared to less specific CNV callers, resulting in a much higher positive predictive value (precision). MDTS also exhibited much better scalability, and is available as open source software at github.com/JMF47/MDTS.

bioinformatics

Genotype Imputation Performance of Three Reference Panels Using African Ancestry Individuals

Genotype imputation is used to estimate unobserved genotypes from genome-wide maker data, to increase genome coverage and power for genome-wide association studies. Imputation has been most successful for European ancestry populations in which very large reference panels are available. Smaller subsets of African descent populations are available in 1000 Genomes (1000G), the Consortium on Asthma among African-Ancestry Populations in the Americas (CAAPA) and the Haplotype Reference Consortium (HRC). We aimed to compare the performance of these reference panels when imputing variation in 3,747 African Americans (AA) from 2 cohorts (HCV and COPDGene) genotyped using the Illumina Omni family of microarrays. The haplotypes of 2,504 individuals (from 1000G), 883 (from CAAPA) and 32,611 (from HRC) were used as reference. We compared the performance of these panels based on number of variants, imputation quality, imputation accuracy and coverage. In both cohorts, 1000G imputed 1.5-1.6x more variants compared to CAAPA and 1.2x more variants than HRC. Similar findings were observed for variants with higher imputation quality (R2>0.5) and for rare, low frequency, and common variants. When merging the results of the three panels the total number of imputed variants was 62M-63M with 20M overlapping variants imputed by all three panels, and a range of 5 to 15M unique variants imputed exclusively with one of the three panels. For overlapping variants, imputation quality was highest for HRC, followed by 1000G, then CAAPA, and improved as the minor allele frequency increased. The 1000G, HRC and CAAPA participants of African ancestry provided high performance and accuracy for imputation of African American admixed individuals, increasing the total number of variants with high quality available for subsequent analyses. These three panels are complementary and would benefit from the development of an integrated African reference panel, including data from multiple sources and populations.

genetics

Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome

A primary goal of The Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to develop an African Diaspora Power Chip (ADPC), a genotyping array consisting of tagging SNPs, useful in comprehensively identifying African specific genetic variation. This array is designed based on the novel variation identified in 642 CAAPA samples of African ancestry with high coverage whole genome sequence data (~30x depth). This novel variation extends the pattern of variation catalogued in the 1000 Genomes and Exome Sequencing Projects to a spectrum of populations representing the wide range of West African genomic diversity. These individuals from CAAPA also comprise a large swath of the African Diaspora population and incorporate historical genetic diversity covering nearly the entire Atlantic coast of the Americas. Here we show the results of designing and producing such a microchip array. This novel array covers African specific variation far better than other commercially available arrays, and will enable better GWAS analyses for researchers with individuals of African descent in their study populations. A recent study1 cataloging variation in continental African populations suggests this type of African-specific genotyping array is both necessary and valuable for facilitating large-scale GWAS in populations of African ancestry.

genomics