bioRxiv Science⌕ Search

Biology subjects

Omae, K.

Publications and source records attributed to Omae, K..

6 recordsLinked to original sources

Functional Unknomics of the SAR11 clade using bioinformatics approaches

A substantial fraction of the genes in marine bacteria lack detectable sequence similarity to genes with known functions. These functionally uncharacterized genes--collectively referred to as the "unknome"--represent a largely unexplored genetic repertoire harboring insights into marine bacterial ecology. In this study, we explored the function of the unknome of SAR11 clade, the most abundant bacterial lineage in the ocean, with a particular focus on genes that provide insight into their ecology. Based on the COG classification, approximately 58% of SAR11 ortholog groups were classified as unknome. Among the SAR11 unknome, we successfully inferred the functions of 69 ortholog groups that are conserved in SAR11 clade by protein structure similarity searches and genomic context analyses. These ortholog groups include putative transporter components, supporting the current ecological understanding that SAR11 clade is specialized in substrate uptake to adapt to oligotrophic marine environments. Furthermore, structural analysis indicated that the DUF2237-containing protein, enriched in marine environments, has potential interactions with purine nucleotide-containing compounds. This may suggest the existence of unique nucleotide utilization mechanisms in marine bacteria. In addition, we found candidates of virus defense systems within the unknome, which demonstrates that diverse defense systems are present at least in one-third of the cultured SAR11 strains. The conservation of these viral defense systems, even within streamlined SAR11 genomes, suggests that they confer significant ecological advantages. Our exploration provided insights into the genetic basis of bottom-up processes (adaptation to oligotrophic environments) and top-down processes (antiviral defense strategy) contributing to ecological success of SAR11.

microbiology↗

The evolution lifecycle of ribosome hibernation factors

Bacteria defend against hostile environments through a variety of molecular mechanisms, including ribosome hibernation. Previously, bacteria were shown to initiate ribosome hibernation by activating protective proteins known as hibernation factors. It was demonstrated that hibernation factors prevent ribosome degradation by nucleases, which allows bacteria to safely store their inactive ribosomes and survive under starvation or persistent stress. Because homologs of hibernation factors were found in diverse lineages of bacteria, it is currently assumed that the mechanism of ribosome hibernation is highly conserved across species. Here, we assess 46,015 complete bacterial genomes to reveal the principles underlying the origin and evolution of these essential proteins in bacterial cells. We find that hibernation factors emerged in ancient bacteria as relatively large proteins that then gradually reduced in size and have undergone complete extinction in over 10% of studied bacteria. We then demonstrate that the degeneration of ancient hibernation factors is often accompanied by "borrowing" hibernation factors from other species via gene transfers, de novo gene birth or fusion of truncated hibernation factors with fragments from other stress-response proteins. These findings reveal a unique evolutionary pathway in which bacteria respond to the reductive evolution of hibernation machinery by inventing novel hibernation mechanisms, thus restoring their capacity to survive starvation and stress. This model implies that most ribosome hibernation factors are yet to be discovered and predicts the organisms that rely on currently unknown hibernation mechanisms.

microbiology↗

A Novel Insertion Site for Group I Introns in tRNA Genes of Patescibacteria

Patescibacteria is a bacterial phylum with small genomes, frequent loss of essential genes, and the presence of introns. While many aspects of Patescibacteria remain enigmatic, an intriguing feature is the widespread occurrence of introns within their compact genomes. To better understand the diversity, roles, and evolution of bacterial introns, we focused on tRNA introns and analyzed Patescibacteria complete genome. Notably, 20% of these genomes lacked at least one tRNA gene for a canonical amino acid, primarily tRNAAsn and tRNAAsp, whereas other tRNA genes were readily detected. This observation led us to conduct further analyses, resulting in the discovery of a novel group I intron insertion site at position 35/36 within the anticodon loop that likely prevented detection by conventional annotation tools. Splicing assays demonstrated that these bacterial introns are catalytically active and capable of self-splicing. To assess the broader distribution of this insertion site across bacteria, we analyzed 4,934 bacterial genomes and identified 269 group I introns within tRNA genes across 14 phyla. Nearly 70% of introns at position 35/36 originate from Patescibacteria, indicating that this feature is largely confined to the phylum. Subgroup classification showed that 79% of all tRNA introns belonged to the IC subgroup, whereas almost all Patescibacteria introns were assigned to IA, suggesting a distinct evolutionary origin. As most tRNA introns lacked homing endonuclease genes, horizontal transfer appears limited. Collectively, these findings advance our understanding of the phylogenetic distribution and evolutionary history of bacterial group I introns in tRNAs, with particular emphasis on Patescibacteria. IMPORTANCEGroup I introns in bacterial tRNA genes were previously known only in a limited number of phyla. Our study expands this knowledge by identifying a novel insertion position in tRNA genes of phylum Patescibacteria and mapping their phylogenetic distribution across bacterial lineages. Our result revealed that group I introns inserted in tRNA genes differed in subgroups between Patescibacteria and other bacteria, highlighting the evolutionary uniqueness of introns of Patescibacteria. Additionally, we found that group I introns are maintained in 43% of bacterial phyla, with tRNA insertions being the most common. Our findings highlight that even in complete genomes, the presence of group I introns can hinder the detection of all 20 canonical tRNA genes by conventional tRNA annotation tools. This study illustrates the overlooked phylogenetic distribution of group I introns across the bacterial domain.

microbiology↗

Bioinformatics classification of the MgtE Mg2+ channel and de novo protein design for the stabilization of its novel subclass

MgtE channels play crucial roles in Mg{superscript 2} homeostasis and are implicated in bacterial survival under antibiotic exposure. Previous structural and biophysical studies have predominantly focused on the Thermus thermophilus MgtE, leaving the structural and mechanistic diversity of MgtE family proteins largely unexplored. In this study, using a genome mining approach, we identified diverse MgtE homologs, including a novel subclass termed the "mini-N type," which lacks the canonical cytoplasmic N and CBS domains but possesses a unique small N-like domain. Despite extensive expression screening, mini-N type homologs could not be stably purified. To address this issue, we designed a series of de novo proteins and determined their crystal structures. A selected de novo protein was fused to a mini-N type MgtE, enabling successful purification and preliminary cryo-EM imaging. Our findings demonstrate that de novo designed protein fusions can serve as powerful tools for stabilizing and purifying otherwise unstable membrane proteins, opening new avenues for the structural and functional studies of otherwise inaccessible membrane proteins.

biophysics↗

CORGIAS: identifying correlated gene pairs by considering evolutionary history in a large-scale prokaryotic genome dataset

The recent expansion of prokaryotic genomes reveals many ortholog groups (OGs) whose function cannot be inferred from conventional, sequence similarity-based annotation methods, especially in metagenome-assembled genomes. Phylogenetic profiling is one of the promising methods to annotate these OGs, by identifying functional relationships of OGs using co- or anti-occurrence of OGs distributions, not sequence similarity. Here, we proposed two new phylogenetic methods for large-scale data, Ancestral State Adjustment (ASA) and Simultaneous EVolution test (SEV), which consider the ancestral state of gene presence/absence. In evaluations using three distinct prokaryotic datasets, ASA and SEV showed better or comparable performance to both established and recently proposed methods for large-scale data. We compared the functionally related genes detected by each method and found that SEV and its predecessor can identify slowly evolving genes, such as housekeeping genes. In contrast, ASA and its predecessors can detect functionally related genes that tend to be gained or lost in a fixed-order, indicating a strong evolutionary constraint that provides clues for functional prediction. Using matrix multiplication, we showed that SEV is scalable in the latest genome databases.

bioinformatics↗

Repeatability of protein structural evolution following convergent gene fusions

Convergent evolution of proteins provides insights into repeatability of genetic adaptation. While local convergence of proteins at residue or domain level has been characterized, global structural convergence by inter-domain/molecular interactions remains largely unknown. Here we present structural convergent evolution on fusion enzymes of aldehyde dehydrogenases (ALDHs) and alcohol dehydrogenases (ADHs). We discovered BdhE (bifunctional dehydrogenase E), an enzyme clade that emerged independently from the previously known AdhE family through distinct gene fusion events. AdhE and BdhE showed shared enzymatic activities and non-overlapping phylogenetic distribution, suggesting common functions in different species. Cryo-electron microscopy revealed BdhEs form donut-like homotetramers, contrasting AdhEs helical homopolymers. Intriguingly, despite distinct quaternary structures and >70% unshared amino acids, both enzymes form resembled dimeric structure units by ALDH-ADH interactions via convergently elongated loop structures. These findings suggest convergent gene fusions recurrently led to substrate channeling evolution to enhance two-step reaction efficiency. Our study unveils structural convergence at inter-domain/molecular level, expanding our knowledges on patterns behind molecular evolution exploring protein structural universe.

evolutionary biology↗