bioRxiv ScienceSearch

Biology subjects

Lage, K.

Publications and source records attributed to Lage, K..

7 recordsLinked to original sources

Components of genetic associations across 2,138 phenotypes in the UK Biobank highlight novel adipocyte biology

To characterize latent components of genetic associations, we applied truncated singular value decomposition (DeGAs) to matrices of summary statistics derived from genome-wide association analyses across 2,138 phenotypes measured in 337,199 White British individuals in the UK Biobank study. We systematically identified key components of genetic associations and the contributions of variants, genes, and phenotypes to each component. As an illustration of the utility of the approach to inform downstream experiments, we report putative loss of function variants, rs114285050 (GPR151) and rs150090666 (PDE3B), that substantially contribute to obesity-related traits, and experimentally demonstrate the role of these genes in adipocyte biology. Our approach to dissect components of genetic associations across human phenotypes will accelerate biomedical hypothesis generation by providing insights on previously unexplored latent structures.

genetics

Developing a network view of type 2 diabetes risk pathways through integration of genetic, genomic and functional data

Genome wide association studies (GWAS) have identified several hundred susceptibility loci for Type 2 Diabetes (T2D). One critical, but unresolved, issue concerns the extent to which the mechanisms through which these diverse signals influencing T2D predisposition converge on a limited set of biological processes. However, the causal variants identified by GWAS mostly fall into non-coding sequence, complicating the task of defining the effector transcripts through which they operate. Here, we describe implementation of an analytical pipeline to address this question. First, we integrate multiple sources of genetic, genomic, and biological data to assign positional candidacy scores to the genes that map to T2D GWAS signals. Second, we introduce genes with high scores as seeds within a network optimization algorithm (the asymmetric prize-collecting Steiner Tree approach) which uses external, experimentally-confirmed protein-protein interaction (PPI) data to generate high confidence subnetworks. Third, we use GWAS data to test the T2D-association enrichment of the \"non-seed\" proteins introduced into the network, as a measure of the overall functional connectivity of the network. We find: (a) non-seed proteins in the T2D protein-interaction network so generated (comprising 705 nodes) are enriched for association to T2D (p=0.0014) but not control traits; (b) stronger T2D-enrichment for islets than other tissues when we use RNA expression data to generate tissue-specific PPI networks; and (c) enhanced enrichment (p=3.9xl0-5) when we combine analysis of the islet-specific PPI network with a focus on the subset of T2D GWAS loci which act through defective insulin secretion. These analyses reveal a pattern of non-random functional connectivity between causal candidate genes atT2D GWAS loci, and highlight the products of genes including YWHAG, SMAD4 or CDK2 as contributors to T2D-relevant islet dysfunction. The approach we describe can be applied to other complex genetic and genomic data sets, facilitating integration of diverse data types into disease-associated networks.\n\nAuthor summaryWe were interested in the following question: as we discover more and more genetic variants associated with a complex disease, such as type 2 diabetes, will the biological pathways implicated by those variants proliferate, or will the biology converge onto a more limited set of aetiological processes? To address this, we first took the 1895 genes that map to ~100 type 2 diabetes association signals, and pruned these to a set of 451 for which combined genetic, genomic and biological evidence assigned the strongest candidacy with respect to type 2 diabetes pathogenesis. We then sought to maximally connect these genes within a curated protein-protein interaction network. We found that proteins brought into the resulting diabetes interaction network were themselves enriched for diabetes association signals as compared to appropriate control proteins. Furthermore, when we used tissue-specific RNA abundance data to filter the generic protein-protein network, we found that the enrichment for type 2 diabetes association signals was enhanced within a network filtered for pancreatic islet expression, particularly when we selected the subset of diabetes association signals acting through reduced insulin secretion. Our data demonstrate convergence of the biological processes involved in type 2 diabetes pathogenesis and highlight novel contributors.

bioinformatics

Open Community Challenge Reveals Molecular Network Modules with Key Roles in Diseases

Identification of modules in molecular networks is at the core of many current analysis methods in biomedical research. However, how well different approaches identify disease-relevant modules in different types of gene and protein networks remains poorly understood. We launched the "Disease Module Identification DREAM Challenge", an open competition to comprehensively assess module identification methods across diverse protein-protein interaction, signaling, gene co-expression, homology, and cancer-gene networks. Predicted network modules were tested for association with complex traits and diseases using a unique collection of 180 genome-wide association studies (GWAS). Our critical assessment of 75 contributed module identification methods reveals novel top-performing algorithms, which recover complementary trait-associated modules. We find that most of these modules correspond to core disease-relevant pathways, which often comprise therapeutic targets and correctly prioritize candidate disease genes. This community challenge establishes benchmarks, tools and guidelines for molecular network analysis to study human disease biology (https://synapse.org/modulechallenge).

bioinformatics

Detecting cancer vulnerabilities through gene networks under purifying selection in 4,700 cancer genomes

Large-scale cancer sequencing studies have uncovered dozens of mutations critical to cancer initiation and progression. However, a significant proportion of genes linked to tumor propagation remain hidden, often due to noise in sequencing data confounding low frequency alterations. Further, genes in networks under purifying selection (NPS), or those that are mutated in cancers less frequently than would be expected by chance, may play crucial roles in sustaining cancers but have largely been overlooked. We describe here a statistical framework that identifies genes that have a first order protein interaction network significantly depleted for mutations, to elucidate key genetic contributors to cancers. Not reliant on and thus, unbiased by, the gene of interests mutation rate, our approach has identified 685 putative genes linked to cancer development. Comparative analysis indicates statistically significant enrichment of NPS genes in previously validated cancer vulnerability gene sets, while further identifying novel cancer-specific candidate gene targets. As more tumor genomes are sequenced, integrating systems level mutation data through this network approach should become increasingly useful in pinpointing gene targets for cancer diagnosis and treatment.

bioinformatics

A unified web platform for network-based analyses of genomic data

Functional genomics networks are widely used to identify unexpected pathway relationships in large genomic datasets. However, it is challenging to quantitatively compare the signal-to-noise ratio of different networks, the biology they describe, and to identify the optimal network to interpret a particular genetic dataset. Via GeNets users can train a machine-learning model (Quack) to make such comparisons; and they can execute, store, and share analyses of genetic and RNA sequencing datasets.

genomics

Expanding discovery from cancer genomes by integrating protein network analyses with in vivo tumorigenesis assays

Approaches that integrate molecular network information and tumor genome data could complement gene-based statistical tests to identify likely new cancer genes, but are challenging to validate at scale and their predictive value remains unclear. We developed a robust statistic (NetSig) that integrates protein interaction networks and data from 4,742 tumor exomes and used it to accurately classify known driver genes in 60% of tested tumor types and to predict 62 new candidates. We designed a quantitative experimental framework to compare the in vivo tumorigenic potential of NetSig candidates, known oncogenes and random genes in mice showing that NetSig candidates induce tumors at rates comparable to known oncogenes and 10-fold higher than random genes. By reanalyzing nine tumor-inducing NetSig candidates in 242 patients with oncogene-negative lung adenocarcinomas, we find that two (AKT2 and TFDP2) are significantly amplified. Overall, we illustrate a scalable integrated computational and experimental workflow to expand discovery from cancer genomes.

cancer biology

Genoppi: A web application for interactive integration of experimental proteomics results with genetic datasets

SummaryIntegrating protein-protein interaction experiments and genetic datasets can lead to new insight into the cellular processes implicated in diseases, but this integration is technically challenging. Here, we present Genoppi, a web application that integrates quantitative interaction proteomics data and results from genome-wide association studies or exome sequencing projects, to highlight biological relationships that might otherwise be difficult to discern. Written in R, Python and Bash script, Genoppi is a user-friendly framework easily deployed across Mac OS and Linux distributions.\n\nAvailabilityGenoppi is open source and available at https://github.com/lagelab/Genoppi\n\nContactaprilkim@broadinstitute.org and lage.kasper@mgh.harvard.edu

bioinformatics