bioRxiv Science⌕ Search

Biology subjects

Cole, B.

Publications and source records attributed to Cole, B..

6 recordsLinked to original sources

ECLIPSER: identifying causal cell types and genes for complex traits through single cell enrichment of e/sQTL-mapped genes in GWAS loci

SummaryECLIPSER was developed to identify pathogenic cell types and cell type-specific genes that may affect complex disease susceptibility and trait variation by integrating single cell data with known GWAS loci. ECLIPSER maps genes to GWAS loci for a given complex trait based on expression and splicing quantitative trait loci (e/sQTLs) and other functional data, and tests whether the mapped genes are enriched for cell type-specific expression in particular cell types using single-cell/nucleus RNA-seq data from one or more tissues of interest. A Bayesian Fishers exact test is used to compute fold-enrichment significance. We demonstrate the application of ECLIPSER on various skin diseases and traits using snRNA-seq of healthy human skin samples. Availability and ImplementationThe source code and documentation for ECLIPSER and a Jupyter notebook for generating output tables and figures are available at https://github.com/segrelabgenomics/ECLIPSER. The source code for GWASvar2gene that maps genes to GWAS loci based on e/sQTLs is available at https://github.com/segrelabgenomics/GWASvar2gene. The analysis presented here used data from GTEx (https://gtexportal.org/home/datasets) and Open Targets Genetics (https://genetics-docs.opentargets.org/data-access/graphql-api), but can also be applied to other GWAS variant lists and QTL studies. Data used to reproduce the results of the paper are available in Supplementary data.

bioinformatics↗

LoFTK: a framework for fully automated calculation of predicted Loss-of-Function variants

MotivationLoss-of-Function (LoF) variants in human genes are important due to their impact on clinical phenotypes and frequent occurrence in the genomes of healthy individuals. Current approaches predict high-confidence LoF variants without identifying the specific genes or the number of copies they affect. Moreover, there is a lack of methods for detecting knockout genes caused by compound heterozygous (CH) LoF variants. ResultsWe have developed the Loss-of-Function ToolKit (LoFTK), which allows efficient and automated prediction of LoF variants from both genotyped and sequenced genomes. LoFTK enables the identification of genes that are inactive in one or two copies and provides summary statistics for downstream analyses. LoFTK can identify CH LoF variants, which result in LoF genes with two copies lost. Using data from parents and offspring we show that 96% of CH LoF genes predicted by LoFTK in the offspring have the respective alleles donated by each parent. Availability and implementationLoFTK is an open source software and is freely available to non-commercial users at https://github.com/CirculatoryHealth/LoFTK Contactj.vansetten@umcutrecht.nl Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Plant Metabolic Network: A multi-species resource of plant metabolic information

Plant metabolism is a pillar of our ecosystem, food security, and economy. To understand and engineer plant metabolism, we first need a comprehensive and accurate annotation of all metabolic information across plant species. As a step towards this goal, we previously created the Plant Metabolic Network (PMN), an online resource of curated and computationally predicted information about the enzymes, compounds, reactions, and pathways that make up plant metabolism. Here we report PMN 15, which contains genome-scale metabolic pathway databases of 126 algal and plant genomes, ranging from model organisms to crops to medicinal plants, and new tools for analyzing and viewing metabolism information across species and integrating omics data in a metabolic context. We systematically evaluated the quality of the databases, which revealed that our semi-automated validation pipeline dramatically improves the quality. We then compared the metabolic content across the 126 organisms using multiple correspondence analysis and found that Brassicaceae, Poaceae, and Chlorophyta appeared as metabolically distinct groups. To demonstrate the utility of this resource, we used recently published sorghum transcriptomics data to discover previously unreported trends of metabolism underlying drought tolerance. We also used single-cell transcriptomics data from the Arabidopsis root to infer cell-type specific metabolic pathways. This work shows the continued growth and refinement of the PMN resource and demonstrates its wide-ranging utility in integrating metabolism with other areas of plant biology. One-sentence SummaryThe Plant Metabolic Network is a collection of databases containing experimentally-supported and predicted information about plant metabolism spanning many species.

plant biology↗

In-depth single-cell analysis of translation-competent HIV-1 reservoirs identifies cellular sources of plasma viremia

Clonal expansion of HIV-infected cells contributes to the long-term persistence of the HIV reservoir in ART-suppressed individuals. However, the contribution to plasma viremia from cell clones that harbor inducible proviruses is poorly understood. Here, we describe a single-cell approach to simultaneously sequence the TCR, integration sites and proviral genomes from translation-competent reservoir cells, called STIP-Seq. By applying this approach to blood samples from eight participants, we showed that the translation-competent reservoir mainly consists of proviruses with short deletions at the 5-end of the genome, often involving the major splice donor site. TCR and integration site sequencing revealed that antigen-responsive cells can harbor inducible proviruses integrated into cancer-related genes. Furthermore, we found several matches between proviruses retrieved with STIP-Seq and plasma viruses obtained during ART and upon treatment interruption, showing that STIP-Seq can capture clones that are responsible for low-level viremia or viral rebound.

microbiology↗

In-depth characterization of HIV-1 reservoirs reveals links to viral rebound during treatment interruption

The HIV-1 reservoir is composed of cells harboring latent proviruses that are capable of contributing to viremia upon antiretroviral treatment (ART) interruption. Although this reservoir is known to be maintained by clonal expansion, the contribution of large, infected cell clones to residual viremia and viral rebound remains underexplored. Here, we conducted an extensive analysis on four ART-treated individuals who underwent an analytical treatment interruption (ATI). We performed subgenomic (V1-V3 env), near full-length proviral and integration site sequencing, and used multiple displacement amplification to sequence both the integration site and provirus from single HIV-infected cells. We found eight proviruses that could phylogenetically be linked to plasma virus obtained before or during the ATI. This study highlights a role for HIV-infected cell clones in the maintenance of the replication-competent reservoir and suggests that infected cell clones can directly contribute to rebound viremia upon ATI.

microbiology↗

Integration of Protein Interactome Networks with Congenital Heart Disease Variants Reveals Candidate Disease Genes

Congenital heart disease (CHD) is present in 1% of live births, yet identification of causal mutations remains a challenge despite large-scale genomic sequencing efforts. We hypothesized that genetic determinants for CHDs may lie in protein interactomes of GATA4 and TBX5, two transcription factors that cause CHDs. Defining their interactomes in human cardiac progenitors via affinity purification-mass spectrometry and integrating results with genetic data from the Pediatric Cardiac Genomic Consortium revealed an enrichment of de novo variants among proteins that interact with GATA4 or TBX5. A consolidative score that prioritized interactome members based on variant, gene, and proband features identified likely CHD-causing genes, including the epigenetic reader GLYR1. GLYR1 and GATA4 widely co-occupied cardiac developmental genes, resulting in co-activation, and the GLYR1 missense variant associated with CHD disrupted interaction with GATA4. This integrative proteomic and genetic approach provides a framework for prioritizing and interrogating the contribution of genetic variants in disease.

genetics↗