bioRxiv ScienceSearch

bioRxiv · 10.64898/2026.08.29.748005

Evolutionary origins of protein novelty across an entire yeast subphylum

Abstract

Novel protein-coding sequences fuel molecular and cellular evolutionary innovations and frequently contribute to species-defining characteristics. They can originate either de novo from previously noncoding sequences or through extreme divergence of already coding ones. How frequently each mechanism occurs and how they shape the structural and functional potential of the resulting proteins remains unclear. Here, we conducted a broad computational investigation of genetic and protein novelty throughout the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 6,000 novel genes resulting from extreme sequence divergence, revealing two quantitatively similar but qualitatively distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 20% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that ''invention'' of new folds occurs more frequently through de novo emergence. We also illustrate cases of evolutionary ''re-invention'' of existing protein folds from noncoding sequences. Our work deepens our understanding of the origins and importance of novel proteins, opening new directions for further structural and functional characterization.

Explore related subjects

Keep this discovery

BibTeXRIS

Tassios, E., Pyrgelis, N., Rinker, D., Tzermpou, E. M., Hittinger, C. T., Rokas, A., Nikolaou, C., Vakirlis, N.. 2026-09-01. Evolutionary origins of protein novelty across an entire yeast subphylum. https://doi.org/10.64898/2026.08.29.748005

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

Dog-wise canine gut metagenome assemblies with reconstructed bacterial genomes and viral candidates

Long-read metagenomic sequencing can improve genome recovery from complex gut microbial communities, yet directly reusable canine gut genome resources remain limited. Here we describe DogMAG, a canine gut metagenome resource based on dog-wise long-read and hybrid assemblies generated by grouping sequencing libraries according to canonical dog identity before assembly. The final dataset comprises 41 assemblies linked to 277 FASTQ records, including 30 Flye long-read-only and 11 OPERA-MS hybrid assemblies. A single integrated BASALT workflow produced 11,276 selected bin/version records, followed by explicit quality-based re-selection of 3,418 medium-quality-or-better metagenome-assembled genome candidates. External dRep dereplication yielded 792 strain-like representatives at 99% average nucleotide identity and 135 species/SGB-like representatives at 95%. GTDB-Tk classified all 792 representatives as Bacteria. Viral screening identified 22,068 geNomad predictions, of which 3,374 Complete, High-quality or Medium-quality viral/proviral candidate rows passed CheckV filtering with contamination [≤]10%. DogMAG provides assemblies, genome and viral candidate sequences, metadata, provenance tables and workflow scripts for reuse, benchmarking and reanalysis.

microbiology

Comparative genomics of clinical isolates of Pseudomonas aeruginosa from cystic fibrosis patients in Mexico

Pseudomonas aeruginosa (P. aeruginosa) is the primary pathogen responsible for morbidity and mortality in patients with cystic fibrosis (CF). Its genomic plasticity and constant selective pressure from antimicrobial treatments have favored the emergence of multidrug-resistant clones. This study conducted a comparative genomic analysis of 41 P. aeruginosa isolated from pediatric patients with CF in Mexico from 2015 to 2024, with the aim of characterizing their evolutionary dynamics, resistome, and virulome. Whole-genome sequencing (MGI, Illumina, and PacBio platforms) was used, with de novo assemblies performed using Unicycler v0.4.8 on the BV-BRC platform. The databases used for the resistome were CARD and NDARO, and for the virulome, VFDB. Phylogenetic reconstruction was based on core-genome alignments generated with Roary v3.13.0, with maximum likelihood reconstruction performed in IQ-TREE v2.1.2. The statistical significance of the segregation of resistance and virulence patterns was evaluated using PERMANOVA analysis. The results revealed a significant clonal prevalence of sequence types (ST) 307 and ST 167. Phylogenomic analysis grouped the isolates into three main clades; Clade 1 stood out for having the highest resistance gene load (mean of 75 genes/genome), establishing itself as the main reservoir of multidrug-resistant profiles. Genotype-phenotype concordance reached 65.5% overall, with high accuracy for aminoglycosides (87.8%) and fluoroquinolones (82.9%). Furthermore, virulome analysis identified 67 distinct patterns that were significantly segregated among the clades (PERMANOVA: R2=0.31, p=0.001). These findings demonstrate that the evolution of P. aeruginosa lineages in the pediatric clinical setting involves parallel and coordinated adaptations in both their resistance potential and their virulence arsenal. This study underscores the need to adopt a multidisciplinary approach to the clinical management of chronic P. aeruginosa infections in pediatric patients. The persistence of extensively drug-resistant (XDR) strains calls for the integration of genomic surveillance and functional diagnostics, as well as the search for therapeutic alternatives for the clinical management of patients with cystic fibrosis.

microbiology

Using sequence-to-function models to interpret archaic hominin introgression

Understanding the functional impact of archaic hominin introgression remains challenging due to the poor representation of global introgression in publicly available genomics resources. Sequence-to-function models can predict the effects of any possible variant in the human genome and may fill this gap. Here, we used AlphaGenome to predict the effects of 144,139 introgressed SNPs segregating in present-day individuals of Papuan genetic ancestry. AlphaGenome's chromatin accessibility predictions recapitulate experimentally observed effects, but gene expression performs no better than chance. Predictions correlate more strongly with an independent reporter assay of single-variant activity than with the same variants' effects in live cells, indicating that AlphaGenome captures the regulatory potential of individual variants more reliably. Predictions carry tissue specificity, allowing us to predict specific tissues potentially impacted by introgressed haplotypes. We identify genes, including JAK1 and TAB2, that are associated with haplotypes that contain an excess of variants predicted by AlphaGenome to have large impacts on chromatin accessibility. Finally, we highlight the challenges and limitations associated with using sequence-to-function models for introgressed variant effect prediction, and show that while AlphaGenome's chromatin accessibility predictions can aid in prioritising candidate functional regions, expression predictions and the assignment of variants to target genes remain as open challenges.

genomics