bioRxiv Science⌕ Search

Biology subjects

Li, D. B.

Publications and source records attributed to Li, D. B..

5 recordsLinked to original sources

Coevolutionary mining of prokaryotic non-coding elements with a genome language model

Microbial genomes encode compact molecular machines and diverse non-coding RNAs (ncRNAs) essential to gene regulation, pathogenesis, and many foundational biotechnologies. However, annotation remains largely protein-centric and homology-driven. Here, we introduce Minerva, a framework for coevolutionary mining that uses genome language models to predict local interactions directly from sequence as two-dimensional maps. Introducing two complementary techniques, categorical Jacobian fingerprinting and interaction heads, we demonstrate fast, accurate, alignment-free prediction of ncRNA base-pairing, monomeric protein contacts, and repetitive sequence motifs. Applied to 150 bacterial genomes, Minerva recovers known systems and predicts that 84.3% of predicted intergenic base-pairing falls outside of known annotations. In Pseudomonas, we find that the widespread TwoAYGGAY ncRNA family carries large secondary-structure extensions and is often flanked by short upstream repetitive motifs and larger downstream genomic repeats. Interpreting the coevolution maps, we find that Minerva emergently detects open-reading-frame (ORF) signatures at the DNA level despite never being trained to do so. In prophages within these genomes, we discover that Unknown Group 27 (UG27) reverse transcriptase systems encode arrays of structurally conserved yet sequence-diverse ncRNAs that template complementary DNA (cDNA) hairpin products. Together, these results establish coevolutionary mining as a scalable route to genome annotation and biological discovery across the rapidly expanding microbial universe.

bioinformatics↗

Microbial mechanistic requirements for eliciting a topical and intranasal immune response

The skin colonist Staphylococcus epidermidis elicits a potent antibody response that can be redirected against an antigen of interest, but this process relies on genetic engineering.1 Here, by adapting bioorthogonal chemistry methods to conjugate antigens to the cell surface, we make the process of generating a commensal vaccine rapid and efficient. A wide variety of bacteria displaying tetanus toxin fragment C (TTFC) elicit an antibody response when applied to mice topically, indicating that the inductive process is not limited to colonists. Colonization occurs at two different sites, the skin and nostril; by colonizing each with S. epidermidis-TTFC, we show that skin colonization yields a moderate IgG response, while nostril colonization elicits a highly potent systemic IgG response and an exuberant IgA response in the nostrils, lungs, and intestine. Two lines of evidence are consistent with the nasal-associated lymphoid tissue (NALT) as the inductive site for nostril colonization: imaging suggesting robust bacterial translocation, and the induction of commensal-specific B cells following colonization. On the skin, TTFC must be conjugated to live S. epidermidis to elicit an antibody response; in the nostrils, live S. epidermidis-TTFC and S. epidermidis mixed with TTFC are equally potent. Commensal vaccination yields a robust response in pet shop mice, and chemical conjugation facilitates antibody responses to a broad array of antigens, including a whole viral capsid and a rotavirus immunogen. Collectively, these findings provide compelling evidence of a translational path for commensal vaccines.

microbiology↗

Phage terminase recognition by the bacterial immune sensors Avs2 and Upx

Prokaryotes employ diverse defense strategies to detect and halt the progression of phage infection. Multiple defense systems sense phage proteins through direct binding, including antiviral STAND NTPases (Avs), which oligomerize upon target recognition to induce programmed cell death. The widespread Avs2 family was previously shown to detect the large terminase subunit of tailed phages, but the mechanism of terminase sensing was unknown. Here, we determine the structural basis of terminase recognition by Avs2 from Escherichia coli (EcAvs2). A cryo-EM structure at 2.3 [A] resolution reveals that EcAvs2 forms a flat, C4-symmetric tetramer in which each protomer is bound to a single terminase monomer. Terminase recognition is mediated by a large, shape complementary binding pocket in the EcAvs2 sensor domain, including specific contacts with an unexpected ATP molecule at the interface of EcAvs2 and terminase. Furthermore, we demonstrate that the defense protein Upx also recognizes diverse phage terminases, despite lacking sequence and structural homology to Avs. AlphaFold 3 models indicate that Upx binds an unfolded state of the core terminase ATPase domain, mediated by {beta}-augmentation. These findings highlight the distinct modes of terminase recognition across structurally diverse defense proteins.

molecular biology↗

Generative design of novel bacteriophages with genome language models

Many important biological functions arise not from single genes, but from complex interactions encoded by entire genomes. Genome language models have emerged as a promising strategy for designing biological systems, but their ability to generate functional sequences at the scale of whole genomes has remained untested. Here, we report the first generative design of viable bacteriophage genomes. We leveraged frontier genome language models, Evo 1 and Evo 2, to generate whole-genome sequences with realistic genetic architectures and desirable host tropism, using the lytic phage {Phi}X174 as our design template. Experimental testing of AI-generated genomes yielded 16 viable phages with substantial evolutionary novelty. Cryo-electron microscopy revealed that one of the generated phages utilizes an evolutionarily distant DNA packaging protein within its capsid. Multiple phages demonstrate higher fitness than {Phi}X174 in growth competitions and in their lysis kinetics. A cocktail of the generated phages rapidly overcomes {Phi}X174-resistance in three E. coli strains, demonstrating the potential utility of our approach for designing phage therapies against rapidly evolving bacterial pathogens. This work provides a blueprint for the design of diverse synthetic bacteriophages and, more broadly, lays a foundation for the generative design of useful living systems at the genome scale.

synthetic biology↗

Genome modeling and design across all domains of life with Evo 2

All of life encodes information with DNA. While tools for sequencing, synthesis, and editing of genomic code have transformed biological research, intelligently composing new biological systems would also require a deep understanding of the immense complexity encoded by genomes. We introduce Evo 2, a biological foundation model trained on 9.3 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life. We train Evo 2 with 7B and 40B parameters to have an unprecedented 1 million token context window with single-nucleotide resolution. Evo 2 learns from DNA sequence alone to accurately predict the functional impacts of genetic variation--from noncoding pathogenic mutations to clinically significant BRCA1 variants--without task-specific finetuning. Applying mechanistic interpretability analyses, we reveal that Evo 2 autonomously learns a breadth of biological features, including exon-intron boundaries, transcription factor binding sites, protein structural elements, and prophage genomic regions. Beyond its predictive capabilities, Evo 2 generates mitochondrial, prokaryotic, and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods. Guiding Evo 2 via inference-time search enables controllable generation of epigenomic structure, for which we demonstrate the first inference-time scaling results in biology. We make Evo 2 fully open, including model parameters, training code, inference code, and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity.

genomics↗