bioRxiv ScienceSearch

Biology subjects

Venter, J. C.

Publications and source records attributed to Venter, J. C..

8 recordsLinked to original sources

Unsupervised integration of multimodal dataset identifies novel signatures of health and disease

Modern medicine is rapidly moving towards a data-driven paradigm based on comprehensive multimodal health assessments. We collected 1,385 data features from diverse modalities, including metabolome, microbiome, genetics and advanced imaging, from 1,253 individuals and from a longitudinal validation cohort of 1,083 individuals. We utilized an ensemble of unsupervised machine learning techniques to identify multimodal biomarker signatures of health and disease risk. In particular, our method identified a set of cardiometabolic biomarkers that goes beyond standard clinical biomarkers, which were used to cluster individuals into distinct health profiles. Cluster membership was a better predictor for diabetes than established clinical biomarkers such as glucose, insulin resistance, and BMI. The novel biomarkers in the diabetes signature included 1-stearoyl-2-dihomo-linolenoyl-GPC and 1-(1-enyl-palmitoyl)-2-oleoyl-GPC. Another metabolite, cinnamoylglycine, was identified as a potential biomarker for both gut microbiome health and lean mass percentage. We also identified an early disease signature for hypertension, and individuals at-risk for a poor metabolic health outcome. We found novel associations between an uremic toxin, p-cresol sulfate, and the abundance of the microbiome genera Intestinimonas and an unclassified genus in the Erysipelotrichaceae family. Our methodology and results demonstrate the potential of multimodal data integration, from the identification of novel biomarker signatures to a data-driven stratification of individuals into disease subtypes and stages -- an essential step towards personalized, preventative health risk assessment.

bioinformatics

Profound perturbation of the human metabolome by obesity

Obesity is a heterogeneous phenotype that is crudely measured by body mass index (BMI). More precise phenotyping and categorization of risk in large numbers of people with obesity is needed to advance clinical care and drug development. Here, we used non-targeted metabolome analysis and whole genome sequencing to identify metabolic and genetic signatures of obesity. We collected anthropomorphic and metabolic measurements at three timepoints over a median of 13 years in 1,969 adult twins of European ancestry and at a single timepoint in 427 unrelated volunteers. We observe that obesity results in a profound perturbation of the metabolome; nearly a third of the assayed metabolites are associated with changes in BMI. A metabolome signature identifies the healthy obese and also identifies lean individuals with abnormal metabolomes - these groups differ in health outcomes and underlying genetic risk. Because metabolome profiling identifies clinically meaningful heterogeneity in obesity, this approach could help select patients for clinical trials.

genetics

No major flaws in "Identification of individuals by trait prediction using whole-genome sequencing data"

In a recently published PNAS article, we studied the identifiability of genomic samples using machine learning methods [Lippert et al., 2017]. In a response, Erlich [2017] argued that our work contained major flaws. The main technical critique of Erlich [2017] builds on a simulation experiment that shows that our proposed algorithm, which uses only a genomic sample for identification, performed no better than a strategy that uses demographic variables. Below, we show why this comparison is misleading and provide a detailed discussion of the key critical points in our analyses that have been brought up in Erlich [2017] and in the media. Further, not only faces may be derived from DNA, but a wide range of phenotypes and demographic variables. In this light, the main contribution of Lippert et al. [2017] is an algorithm that identifies genomes of individuals by combining multiple DNA-based predictive models for a myriad of traits.

genomics

Functional characterization of 3D-protein structures informed by human genetic diversity

Sequence variation data of the human proteome can be used to analyze 3-dimensional (3D) protein structures to derive functional insights. We used genetic variant data from nearly 150,000 individuals to analyze 3D positional conservation in 4,390 protein structures using 481,708 missense and 264,257 synonymous variants. Sixty percent of protein structures harbor at least one intolerant 3D site as defined by significant depletion of observed over expected missense variation. We established an Angstrom-scale distribution of annotated pathogenic missense variants and showed that they accumulate in proximity to the most intolerant 3D sites. Structural intolerance data correlated with experimental functional read-outs in vitro. The 3D structural intolerance analysis revealed characteristic features of ligand binding pockets, orthosteric and allosteric sites. The identification of novel functional 3D sites based on human genetic data helps to validate, rank or predict drug target binding sites in vivo.

genomics

Microbial Metagenome Of Urinary Tract Infection

Urine culture and microscopy techniques are used to profile the bacterial species present in urinary tract infections. To gain insight into the urinary flora in infection and health, we analyzed clinical laboratory features and the microbial metagenome of 121 clean-catch urine samples. 16S rDNA gene signatures were successfully obtained for 116 participants, while whole genome shotgun sequencing data was successfully generated for samples from 49 participants. Analysis of these datasets supports the definition of the patterns of infection and colonization/contamination. Although 16S rDNA sequencing was more sensitive, whole genome shotgun sequencing allowed for a more comprehensive and unbiased representation of the microbial flora, including eukarya and viral pathogens, and of bacterial virulence factors. Urine samples positive by whole genome shotgun sequencing contained a plethora of bacterial (median 41 genera/sample), eukarya (median 2 species/sample) and viral sequences (median 3 viruses/sample). Genomic analyses revealed cases of infection with potential pathogens (e.g., Alloscardovia sp, Actinotignum sp, Ureaplasma sp) that are often missed during routine urine culture due to species specific growth requirements. We also observed gender differences in the microbial metagenome. While conventional microbiological methods are inadequate to identify a large diversity of microbial species that are present in urine, genomic approaches appear to comprehensively and quantitatively describe the urinary microbiome.

microbiology

Precision Medicine Screening Using Whole Genome Sequencing And Advanced Imaging To Identify Disease Risk In Adults

BACKGROUNDProgress in science and technology have created the capabilities and alternatives to symptom-driven medical care. Reducing premature mortality associated with age-related chronic diseases, such as cancer and cardiovascular disease, is an urgent priority we address using advanced screening detection.\n\nMETHODSWe enrolled active adults for early detection of risk for age-related chronic disease associated with premature mortality. Whole genome sequencing together with: global metabolomics, 3D/4D imaging using non-contrast whole body magnetic resonance imaging and echocardiography, and 2-week cardiac monitoring were employed to detect age-related chronic diseases and risk for diseases.\n\nRESULTSWe detected previously unrecognized age-related chronic diseases requiring prompt (<30 days) medical attention in 17 (8%, 1:12) of 209 study participants, including 4 participants with early stage neoplasms (2%, 1:50). Likely mechanistic genomic findings correlating with clinical data were identified in 52 participants (25%. 1:4). More than three-quarters of participants (n=164, 78%, 3:4) had evidence of age-related chronic diseases or associated risk factors.\n\nCONCLUSIONSPrecision medicine screening using genomics with other advanced clinical data among active adults identified unsuspected disease risks for age-related chronic diseases associated with premature mortality. This technology-driven phenotype screening approach has the potential to extend healthy life among active adults through improved early detection and prevention of age-related chronic diseases. Our success provides a scalable strategy to move medical practice and discovery toward risk detection and disease modification thus achieving healthier extension of life.\n\nSIGNIFICANCE STATEMENTAdvances in science and technology have enabled scientists to analyze the human genome cost-effectively and to combine genome sequencing with noninvasive imaging technologies for alternatives to symptom-driven medical care. Using whole genome sequencing and noninvasive 3D/4D imaging technologies we screened 209 adults to detect age-related chronic diseases, such as cancer and cardiovascular disease. We found unrecognized age-related chronic diseases requiring prompt (<30 days) medical attention in 1:12 study participants, likely genomic findings correlating with clinical data in 1:4 participants, and evidence of age-related chronic diseases or associated risk factors in more than 3 of 4 participants. These results demonstrate that genome sequencing with clinical imaging data can be used for screening and early detection of diseases associated with premature mortality.

genomics

Paternally inherited noncoding structural variants contribute to autism

The genetic architecture of autism spectrum disorder (ASD) is known to consist of contributions from gene-disrupting de novo mutations and common variants of modest effect. We hypothesize that the unexplained heritability of ASD also includes rare inherited variants with intermediate effects. We investigated the genome-wide distribution and functional impact of structural variants (SVs) through whole genome analysis ([&ge;]30X coverage) of 3,169 subjects from 829 families affected by ASD. Genes that are intolerant to inactivating variants in the exome aggregation consortium (ExAC) were depleted for SVs in parents, specifically within fetal-brain promoters, UTRs and exons. Rare paternally-inherited SVs that disrupt promoters or UTRs were over-transmitted to probands (P = 0.0013) and not to their typically-developing siblings. Recurrent functional noncoding deletions implicate the gene LEO1 in ASD. Protein-coding SVs were also associated with ASD (P = 0.0025). Our results establish that rare inherited SVs predispose children to ASD, with differing contributions from each parent.

genomics

The human functional genome defined by genetic diversity

Large scale efforts to sequence whole human genomes provide extensive data on the non-coding portion of the genome. We used variation information from 11,257 human genomes to describe the spectrum of sequence conservation in the population. We established the genome-wide variability for each nucleotide in the context of the surrounding sequence in order to identify departure from expectation at the population level (context-dependent conservation). We characterized the population diversity for functional elements in the genome and identified the coordination of conserved sequences of distal and cis enhancers, chromatin marks, promoters, coding and intronic regions. The most context-dependent conserved regions of the genome are associated with unique functional annotations and a genomic organization that spreads up to one megabase. Importantly, these regions are enriched by over 100-fold of non-coding pathogenic variants. This analysis of human genetic diversity thus provides a detailed view of sequence conservation, functional constraint and genomic organization of the human genome. Specifically, it identifies highly conserved non-coding sequences that are not captured by analysis of interspecies conservation and are greatly enriched in disease variants.

genomics