bioRxiv ScienceSearch

Biology subjects

Shiu, S.-H.

Publications and source records attributed to Shiu, S.-H..

7 recordsLinked to original sources

Evolutionary characteristics of intergenic transcribed regions indicate widespread noisy transcription in the Poaceae

Extensive transcriptional activity occurring in unannotated, intergenic regions of genomes has raised the question whether intergenic transcription represents the activity of novel genes or noisy expression. To address this, we evaluated cross-species and post-duplication sequence and expression conservation of intergenic transcribed regions (ITRs) in four Poaceae species. Most ITR sequences are species-specific. Those found across species tend to be more divergent in expression and have more recent duplicates compared to annotated genes. To assess if ITRs are functional (under selection), machine learning models were established in Oryza sativa (rice) that could distinguish between benchmark functional (phenotype genes) and nonfunctional (pseudogenes) sequences with high accuracy based on 44 evolutionary and biochemical features. Based on the prediction models, 584 rice ITRs (8%) are classified as likely functional that tend to have conserved expression and ancient retained duplicates. However, most ITRs do not exhibit sequence or expression conservation across species or following duplication, consistent with computational predictions that suggest 61% ITRs are not under selection. We outline key evolutionary characteristics that are tightly associated with likely-functional ITRs and provide a framework to identify novel genes to improve genome annotation and move toward connecting genotype to phenotype in crop and model systems.

genomics

Predicting cell-cycle expressed genes identifies canonical and non-canonical regulators of time-specific expression in Saccharomyces cerevisiae

The collection all TFs, target genes and their interactions in an organism form a gene regulatory network (GRN), which underly complex patterns of transcription even in unicellular species. However, identifying which interactions regulate expression in a specific temporal context remains a challenging task. With multiple experimental and computational approaches to characterize GRNs, we predicted general and phase-specific cell-cycle expression in Saccharomyces cerevisiae using four regulatory data sets: chromatin immunoprecipitation (ChIP), TF deletion data (Deletion), protein binding microarrays (PBMs), and position weight matrices (PWMs). Our results indicate that the source of regulatory interaction information significantly impacts our ability to predict cell-cycle expression where the best model was constructed by combining selected TF features from ChIP and Deletion data as well as TF-TF interaction features in the form of feed-forward loops. The TFs that were the best predictors of cell-cycle expression were enriched for known cell-cycle regulators but also include regulators not implicated in cell-cycle regulation previously. In addition, ChIP and Deletion datasets led to the identification different subsets of TFs important for predicting cell-cycle expression. Finally, analysis of important TF-TF interaction features suggests that the GRN regulating cell cycle expression is highly interconnected and clustered around four groups of genes, two of which represent known cell-cycle regulatory complexes, while the other two contain TFs that are not known cell-cycle regulators (Ste12-Tex1 and Rap1-Hap1-Msn4), but are nonetheless important to regulating the timing of expression. Thus, not only do our models accurately reflect what is known about the regulation of the S. cerevisiae cell cycle, they can be used to discover regulatory factors which play a role in controlling expression during the cell cycle as well as other contexts with discrete temporal patterns of expression.

genetics

Robust predictions of specialized metabolism genes through machine learning

Plant specialized metabolism (SM) enzymes produce lineage-specific metabolites with important ecological, evolutionary, and biotechnological implications. Using Arabidopsis thaliana as a model, we identified distinguishing characteristics of SM and GM (general metabolism, traditionally referred to as primary metabolism) genes through a detailed study of features including duplication pattern, sequence conservation, transcription, protein domain content, and gene network properties. Analysis of multiple sets of benchmark genes revealed that SM genes tend to be tandemly duplicated, co-expressed with their paralogs, narrowly expressed at lower levels, less conserved, and less well connected in gene networks relative to GM genes. Although the values of each of these features significantly differed between SM and GM genes, any single feature was ineffective at predicting SM from GM genes. Using machine learning methods to integrate all features, a well performing prediction model was established with a true positive rate of 0.87 and a true negative rate of 0.71. In addition, 86% of known SM genes not used to create the machine learning model were predicted as SM genes, further demonstrating its accuracy. We also demonstrated that the model could be further improved when we distinguished between SM, GM, and junction genes responsible for reactions shared by SM and GM pathways. Application of the prediction model led to the identification of 1,217 A. thaliana genes with previously unknown functions, providing a global, high-confidence estimate of SM gene content in a plant genome.\n\nSignificanceSpecialized metabolites are critical for plant-environment interactions, e.g., attracting pollinators or defending against herbivores, and are important sources of plant-based pharmaceuticals. However, it is unclear what proportion of enzyme-encoding genes play roles in specialized metabolism (SM) as opposed to general metabolism (GM) in any plant species. This is because of the diversity of specialized metabolites and the considerable number of incompletely characterized pathways responsible for their production. In addition, SM gene ancestors frequently played roles in GM. We evaluate features distinguishing SM and GM genes and build a computational model that accurately predicts SM genes. Our predictions provide candidates for experimental studies, and our modeling approach can be applied to other species that produce medicinally or industrially useful compounds.

plant biology

Factors influencing gene family size variation among related species in a plant family

Gene duplication and loss contribute to gene content differences as well as phenotypic divergence across species. However, the extent to which gene content varies among closely related plant species and the factors responsible for such variation remain unclear. Here, we used the Solanaceae family as a model to investigate differences in gene family size and the likely factors contributing to these differences. We found that genes in highly variable families have high turnover rate and tend to be involved in processes that have diverged between Solanaceae species, whereas genes in low-variability families tend to have housekeeping roles. In addition, genes in high-and low-variability gene families tend to be duplicated by tandem and whole genome duplication, respectively. This finding together with the observation that genes duplicated by different mechanisms experience different selection pressures suggests that duplication mechanism impacts gene family turnover. We explored using pseudogene number as a proxy for gene loss but discovered that a substantial number of pseudogenes are actually products of pseudogene duplication, contrary to the expectation that most plant pseudogenes are remnants of once-functional duplicates. Our findings reveal complex relationships between variation in gene family size, gene functions, duplication mechanism, and evolutionary rate. The patterns of lineage-specific gene family expansion within the Solanaceae provide the foundation for a better understanding of the genetic basis underlying phenotypic diversity in this economically important family.

evolutionary biology

Defining functional intergenic transcribed regions based on heterogeneous features of phenotype genes and pseudogenes

With advances in transcript profiling, the presence of transcriptional activities in intergenic regions has been well established in multiple model systems. However, whether intergenic expression reflects transcriptional noise or the activity of novel genes remains unclear. We identified intergenic transcribed regions (ITRs) in 15 diverse flowering plant species and found that the amount of intergenic expression correlates with genome size, a pattern that could be expected if intergenic expression is largely non-functional. To further assess the functionality of ITRs, we first built machine learning classifiers using Arabidopsis thaliana as a model that can accurately distinguish functional sequences (phenotype genes) and non-functional ones (pseudogenes and random unexpressed intergenic regions) by integrating 93 biochemical, evolutionary, and sequence-structure features. Next, by applying the models to ITRs, we found that 2,453 (21%) had features significantly similar to phenotype genes and thus were likely parts of functional genes, while an additional 17% resembled benchmark RNA genes. However, [~]60% of ITRs were more similar to nonfunctional sequences and should be considered transcriptional noise unless falsified with experiments. The predictive framework establish here provides not only a comprehensive look at how functional, genic sequences are distinct from likely non-functional ones, but also a new way to differentiate novel genes from genomic regions with noisy transcriptional activities.

genomics

Regulatory Divergence In Wound-Responsive Gene Expression In Domesticated And Wild Tomato

BackgroundThe evolution of cis- and trans-regulatory components of transcription is central to how stress response and tolerance differ across species. However, it remains largely unknown how divergence in TF binding specificity and cis-regulatory sites contribute to the divergence of stress-responsive gene expression between wild and domesticated species.\n\nResultsUsing tomato as model, we analyzed the transcriptional profile of wound-responsive genes in wild Solanum pennellii and domesticated S. lycopersicum. We found that extensive expression divergence of wound-responsive genes is associated with speciation. To assess the degree of trans-regulatory divergence between these two species, 342 and 267 putative cis-regulatory elements (pCREs) in S. lycopersicum and S. pennellii, respectively, were identified that were predictive of wound-induced gene expression. We found that 35-66% of pCREs were conserved across species, suggesting that the remaining proportion (34-65%) of pCREs are species specific. This finding indicates a substantially higher degree of trans-regulatory divergence between these two plant species, which diverged [~]3-7 million years ago, compared to that observed in mouse and human, which diverged [~]100 million years ago. In addition, differences in pCRE sites were significantly associated with differences in wound-responsive gene expression between wild and domesticated tomato orthologs, suggesting the presence of substantial cis-regulatory divergence.\n\nConclusionsOur study provides new insights into the mechanistic basis of how the transcriptional response to wounding is regulated and, importantly, the contribution of cis- and trans-regulatory components to variation in wound-responsive gene expression during species domestication.

evolutionary biology

Asymmetric evolution of the transcription profiles and cis-regulatory sites contributes to the retention of transcription factor duplicates

Transcription factors (TFs) play a key role in regulating plant development and response to environmental stimuli. While most genes revert to single copy after a duplication event, transcription factors are retained at a significantly higher rate. However, it is unclear why TF duplicates have higher rates of retention relative to other genes. In this study, we compared three types of features (expression, sequence, and conservation) of retained TFs following whole genome duplication (WGD) events to genes with other functions, using Arabidopsis thaliana as a model. We found that gene function groups with higher maximum expression but lower mean expression tended to have higher duplicate retention rate post WGD, though TFs in particular are retained more often than would be expected based on the features examined. Conversely, expression of individual genes was not associated with duplication, but sequence conservation was. Furthermore, we found that the evolution of TF expression patterns and cis-regulatory cites favors the partitioning of ancestral states among the resulting duplicates. In particular, we found that one duplicate retains the majority of ancestral expression and cis-regulatory sites, while the \"non-ancestral\" duplicate was enriched for novel regulatory sites. To investigate how this pattern of partitioning pattern evolved, we modeled the retention of ancestral states in duplicate pairs using a system of differential equations. Our findings indicate that duplicate pairs evolve to a partitioned state more often than away from it, which in combination with accumulation of new regulatory sites in non-ancestral duplicates, suggest that selection favors partitioning via neofunctionalization.\n\nAuthor SummaryGene expression is controlled by regulatory proteins known as transcription factors. These factors control how an organism develops and responds to its environment. The evolution of transcription factor functions also contributes to the emergence of new species and crop domestication. In plants, new transcription factors mainly arise due to polyploidy, multiplication of the genome. Although most duplicated copies are lost following a genome duplication event, transcription factors are exceptional because they are often kept. Furthermore, we found that transcription factor duplicates that tend to diverge in how they are expressed and regulated in an unusual way where one copy mirrors the original, pre-duplication functional states of the ancestral gene, while the other loses the ancestral status and instead accumulates novel regulatory sites. Our results suggest these duplicate transcription factors may have been kept because one copy preserve ancestral function while the other has evolved new ones.

evolutionary biology