bioRxiv Science⌕ Search

Biology subjects

Fiamenghi, M. B.

Publications and source records attributed to Fiamenghi, M. B..

5 recordsLinked to original sources

Coevolutionary mining of prokaryotic non-coding elements with a genome language model

Microbial genomes encode compact molecular machines and diverse non-coding RNAs (ncRNAs) essential to gene regulation, pathogenesis, and many foundational biotechnologies. However, annotation remains largely protein-centric and homology-driven. Here, we introduce Minerva, a framework for coevolutionary mining that uses genome language models to predict local interactions directly from sequence as two-dimensional maps. Introducing two complementary techniques, categorical Jacobian fingerprinting and interaction heads, we demonstrate fast, accurate, alignment-free prediction of ncRNA base-pairing, monomeric protein contacts, and repetitive sequence motifs. Applied to 150 bacterial genomes, Minerva recovers known systems and predicts that 84.3% of predicted intergenic base-pairing falls outside of known annotations. In Pseudomonas, we find that the widespread TwoAYGGAY ncRNA family carries large secondary-structure extensions and is often flanked by short upstream repetitive motifs and larger downstream genomic repeats. Interpreting the coevolution maps, we find that Minerva emergently detects open-reading-frame (ORF) signatures at the DNA level despite never being trained to do so. In prophages within these genomes, we discover that Unknown Group 27 (UG27) reverse transcriptase systems encode arrays of structurally conserved yet sequence-diverse ncRNAs that template complementary DNA (cDNA) hairpin products. Together, these results establish coevolutionary mining as a scalable route to genome annotation and biological discovery across the rapidly expanding microbial universe.

bioinformatics↗

Discovery of novel enzybiotic candidates targeting human bacterial pathogens through large-scale viral-host profiling

The rise of antibiotic-resistant bacteria demands alternative therapeutic strategies, with bacteriophage (phage) therapy and phage-derived enzybiotics emerging as promising approaches. However, identifying candidate phages against specific pathogens has historically been a bottleneck due to the need for cultivation methods to assess host range and lytic activity. Advances in metagenomic sequencing and the emergence of large-scale viral genome databases now provide an opportunity to accelerate this process computationally. Here, we present a large-scale mining of the MetaVR database to identify phages targeting human bacterial pathogens. By integrating direct host associations with CRISPR-spacer evidence, we linked 196,472 high-quality and complete viral genomes, representing 42,360 vOTUs, to 618 species of pathogenic and opportunistic bacteria. Functional enrichment analysis revealed distinct genomic signatures with viral lifestyle and host-range breadth: virulent phages were enriched in replication and structural functions, whereas temperate and broad host-range phages were enriched in anti-defense and regulatory modules. To characterize their lytic potential we annotated lysis-related protein families and their structural diversity, identifying 76 structurally novel lysis-associated proteins, including candidates targeting WHO priority pathogens. Focused analysis of endolysins revealed 592 structural clusters, with extensive sharing of endolysin repertoires among ESKAPE pathogens, suggesting candidates for broad-spectrum enzybiotic development. Selection analysis identified 167 endolysin families with sites under positive selection within functional domains, highlighting evolutionary diversification potentially associated with phage-host interactions. Together, our results establish a large-scale framework for connecting human bacterial pathogens to phages and their lytic machinery, providing a resource for prioritizing phage therapy and enzybiotic development.

bioinformatics↗

VPF-Class 2.0: a taxonomy-centered framework for automatic viral classification

Rapid expansion of viral sequence data demands classifiers that scale, track ICTV updates, and provide interpretable evidence. We present VPF-Class 2.0, an updated successor to VPF-Class, centred on the taxonomic classification, that retains marker-driven protein domain detection but replaces rule-based voting with a lightweight supervised model on per-genome marker-composition features. In controlled benchmarks, VPF-Class 2.0 achieves near-perfect family-level performance and strong genus-level accuracy while increasing confident annotation coverage. Under a practical confidence threshold (0.3), performance improves and matches or exceeds representative tools within shared taxonomic scopes. We further introduce an interpretability study that relates errors to the genus specificity of activated markers. Finally, we demonstrate applicability on large real-world viromes with consistent labels and substantial agreement with graph-based classifications. The implementation of VPF-Class 2.0 can be downloaded from https://github.com/luisvidalj/VPFClass2.git.

bioinformatics↗

A genomic atlas of the human gut virome elucidates genetic factors shaping host interactions

Viruses are key modulators of human gut microbiome composition and function. While metagenomic sequencing has enabled culture-independent discovery of gut bacteriophage diversity, existing genomic catalogues suffer from limited geographic representation, sparse taxonomic classification, and insufficient functional annotation, hindering detailed investigation into phage biology. Here, we present the Unified Human Gastrointestinal Virome (UHGV), a collection of 873,994 viral genomes from globally diverse populations that addresses these limitations. UHGV provides high-quality virome references with extensive host predictions, comprehensive functional annotations, protein structures, a classification framework for comparative analysis, and a web portal to facilitate data access. Using UHGV to profile worldwide metagenomes, we found that host range breadth is strongly associated with phage prevalence. Additionally, we identified diversity-generating retroelements and DNA methyltransferases as key factors enabling phage populations to access diverse hosts, revealing how specific genomic features contribute to global phage distribution patterns. UHGV is available at http://uhgv.jgi.doe.gov.

genomics↗

Comparative Genomics of Firmicutes reveals probable adaptations for xylose fermentation in Thermoanaerobacterium saccharolyticum

Second-generation (2G) ethanol is one potential biofuel that could be used to achieve the goal of reducing greenhouse gas emissions. Many challenges still need to be overcome for the feasibility of this technology, most of them related to consumption of xylose, a pentose sugar not easily metabolized by industrial microorganisms. Thus, exploring genes, pathways and other organisms that can ferment xylose is a strategy implemented to solve industrial bottlenecks. Thermoanaerobacterium saccharolyticum (T. sac) is an organism from the firmicutes phylum, capable of naturally fermenting compounds of industrial interest, such as xylan and xylose. Understanding evolutionary adaptations may help not only to solidify this bacterium as a potential substitute to the yeast Saccharomyces cerevisiae in industry, but also bring novel genes and information that can be used for yeast, enhance its fermenting capabilities, and increase production of current bio-platforms. This study presents a deep evolutionary study of members of the firmicutes clade, focusing on adaptations that may be related to overall fermentation metabolism, especially for xylose fermentation. One highlight is the finding of positive selection on a xylose binding protein of the xylFGH operon, close to the annotated sugar binding site, with this protein already being found to be expressed in xylose fermenting conditions in a previous study. Results from this study can serve as basis for searching for candidate genes to use in industrial strains or to improve T. sac as a new microbial cell factory, which may help to solve current problems found in the biofuels industry.

evolutionary biology↗