bioRxiv ScienceSearch

Biology subjects

Wisecaver, J. H.

Publications and source records attributed to Wisecaver, J. H..

4 recordsLinked to original sources

integRATE: a desirability-based data integration framework for the prioritization of candidate genes across heterogeneous omics and its application to preterm birth

BackgroundThe integration of high-quality, genome-wide analyses offers a robust approach to elucidating genetic factors involved in complex human diseases. Even though several methods exist to integrate heterogeneous omics data, most biologists still manually select candidate genes by examining the intersection of lists of candidates stemming from analyses of different types of omics data that have been generated by imposing hard (strict) thresholds on quantitative variables, such as P-values and fold changes, increasing the chance of missing potentially important candidates.\n\nMethodsTo better facilitate the unbiased integration of heterogeneous omics data collected from diverse platforms and samples, we propose a desirability function framework for identifying candidate genes with strong evidence across data types as targets for follow-up functional analysis. Our approach is targeted towards disease systems with sparse, heterogeneous omics data, so we tested it on one such pathology: spontaneous preterm birth (sPTB).\n\nResultsWe developed the software integRATE, which uses desirability functions to rank genes both within and across studies, identifying well-supported candidate genes according to the cumulative weight of biological evidence rather than based on imposition of hard thresholds of key variables. Integrating 10 sPTB omics studies identified both genes in pathways previously suspected to be involved in sPTB as well as novel genes never before linked to this syndrome. integRATE is available as an R package on GitHub (https://github.com/haleyeidem/integRATE).\n\nConclusionsDesirability-based data integration is a solution most applicable in biological research areas where omics data is especially heterogeneous and sparse, allowing for the prioritization of candidate genes that can be used to inform more targeted downstream functional analyses.

genomics

Evidence for loss and adaptive reacquisition of alcoholic fermentation in an early-derived fructophilic yeast lineage

Fructophily is a rare trait that consists in the preference for fructose over other carbon sources. Here we show that in a yeast lineage (the Wickerhamiella/Starmerella, W/S clade) formed by fructophilic species thriving in the floral niche, the acquisition of fructophily is part of a wider process of adaptation of central carbon metabolism to the high sugar environment. Coupling comparative genomics with biochemical and genetic approaches, we show that the alcoholic fermentation pathway was profoundly remodeled in the W/S clade, as genes required for alcoholic fermentation were lost and subsequently re-acquired from bacteria through horizontal gene transfer. We further show that the reinstated fermentative pathway is functional and that an enzyme required for sucrose assimilation is also of bacterial origin, reinforcing the adaptive nature of the genetic novelties identified in the W/S clade. This work shows how even central carbon metabolism can be remodeled by a surge of HGT events.

genomics

Drivers of genetic diversity in secondary metabolic gene clusters in a fungal population

Filamentous fungi produce a diverse array of secondary metabolites (SMs) critical for defense, virulence, and communication. The metabolic pathways that produce SMs are found in contiguous gene clusters in fungal genomes, an atypical arrangement for metabolic pathways in other eukaryotes. Comparative studies of filamentous fungal species have shown that SM gene clusters are often either highly divergent or uniquely present in one or a handful of species, hampering efforts to determine the genetic basis and evolutionary drivers of SM gene cluster divergence. Here we examined SM variation in 66 cosmopolitan strains of a single species, the opportunistic human pathogen Aspergillus fumigatus. Investigation of genome-wide within-species variation revealed five general types of variation in SM gene clusters: non-functional gene polymorphisms, gene gain and loss polymorphisms, whole cluster gain and loss polymorphisms, allelic polymorphisms where different alleles corresponded to distinct, non-homologous clusters, and location polymorphisms in which a cluster was found to differ in its genomic location across strains. These polymorphisms affect the function of representative A. fumigatus SM gene clusters, such as those involved in the production of gliotoxin, fumigaclavine, and helvolic acid, as well as the function of clusters with undefined products. In addition to enabling the identification of polymorphisms whose detection requires extensive genome-wide synteny conservation (e.g., mobile gene clusters and non-homologous cluster alleles), our approach also implicated multiple underlying genetic drivers, including point mutations, recombination, genomic deletion and insertion events, as well as horizontal gene transfer from distant fungi. Finally, most of the variants that we uncover within A. fumigatus have been previously hypothesized to contribute to SM gene cluster diversity across entire fungal classes and phyla. We suggest that the drivers of genetic diversity operating within a fungal species shown here are sufficient to explain SM cluster macroevolutionary patterns.

genomics

A global co-expression network approach for connecting genes to specialized metabolic pathways in plants

Plants produce a tremendous diversity of specialized metabolites (SMs) to interact with and manage their environment. A major challenge hindering efforts to tap this seemingly boundless source of pharmacopeia is the identification of SM pathways and their constituent genes. Given the well-established observation that the genes comprising a SM pathway are co-regulated in response to specific environmental conditions, we hypothesized that genes from a given SM pathway would form tight associations (modules) with each other in gene co-expression networks, facilitating their identification. To evaluate this hypothesis, we used 10 global co-expression datasets--each a meta-analysis of hundreds to thousands of expression experiments--across eight plant model organisms to identify hundreds of modules of co-expressed genes for each species. In support of our hypothesis, 15.3-52.6% of modules contained two or more known SM biosynthetic genes (e.g., cytochrome P450s, terpene synthases, and chalcone synthases), and module genes were enriched in SM functions (e.g., glucoside and flavonoid biosynthesis). Moreover, modules recovered many experimentally validated SM pathways in these plants, including all six known to form biosynthetic gene clusters (BGCs). In contrast, genes predicted based on physical proximity on a chromosome to form plant BGCs were no more co-expressed than the null distribution for neighboring genes. These results not only suggest that most predicted plant BGCs do not represent genuine SM pathways but also argue that BGCs are unlikely to be a hallmark of plant specialized metabolism. We submit that global gene co-expression is a rich, but largely untapped, data source for discovering the genetic basis and architecture of plant natural products, which can be applied even without knowledge of the genome sequence.

plant biology