Coevolutionary mining of prokaryotic non-coding elements with a genome language model
Microbial genomes encode compact molecular machines and diverse non-coding RNAs (ncRNAs) essential to gene regulation, pathogenesis, and many foundational biotechnologies. However, annotation remains largely protein-centric and homology-driven. Here, we introduce Minerva, a framework for coevolutionary mining that uses genome language models to predict local interactions directly from sequence as two-dimensional maps. Introducing two complementary techniques, categorical Jacobian fingerprinting and interaction heads, we demonstrate fast, accurate, alignment-free prediction of ncRNA base-pairing, monomeric protein contacts, and repetitive sequence motifs. Applied to 150 bacterial genomes, Minerva recovers known systems and predicts that 84.3% of predicted intergenic base-pairing falls outside of known annotations. In Pseudomonas, we find that the widespread TwoAYGGAY ncRNA family carries large secondary-structure extensions and is often flanked by short upstream repetitive motifs and larger downstream genomic repeats. Interpreting the coevolution maps, we find that Minerva emergently detects open-reading-frame (ORF) signatures at the DNA level despite never being trained to do so. In prophages within these genomes, we discover that Unknown Group 27 (UG27) reverse transcriptase systems encode arrays of structurally conserved yet sequence-diverse ncRNAs that template complementary DNA (cDNA) hairpin products. Together, these results establish coevolutionary mining as a scalable route to genome annotation and biological discovery across the rapidly expanding microbial universe.