bioRxiv Science⌕ Search

Biology subjects

Maiwald, A.

Publications and source records attributed to Maiwald, A..

3 recordsLinked to original sources

Quantifying evolutionary novelty and design efficiency in generative genome design

Generative genome design models can now produce previously unobserved genome-length sequences, but assessing their capabilities is complicated by limitations in functional prediction. The ability to engineer genomes faster than we can understand them risks creating biosecurity vulnerabilities. To evaluate these potential risks systematically, we propose a framework that distinguishes between (i) evolutionary novelty, quantified through phylogenetic and sequence similarity to natural genomes; and (ii) design efficiency - the efficiency with which a model finds viable sequences compared to simple baseline generators. Applying this framework to bacteriophages designed by the genome language model Evo 2, we find that model likelihood strongly predicts experimental viability, capturing functional constraints beyond simple biological heuristics. However, this efficiency derives largely from staying close to previously observed sequences rather than exploring novel sequence space, reflecting the combined performance of the model and additional filters that were applied to its outputs. Compared to baselines of random mutagenesis and serial passage, the model achieves substantial design efficiency while its outputs remain phylogenetically close to natural genomes. We conclude that the generative capabilities of Evo 2 warrant low to moderate biosecurity concern for de novo hazard creation, although the degree to which these findings generalise to larger or less constrained viral architectures is an open question. Our framework enables an evidence-based capability assessment of generative genome design tools, informing future biosecurity evaluations.

genomics↗

Decode-gLM: Tools to Interpret, Audit, and Steer GenomicLanguage Models

While genomic language models are enabling the de novo design of entire genomes, they remain challenging to interpret, limiting their trustworthiness. Here, we show that sparse autoencoders (SAEs) trained on Nucleotide Transformer activations decompose hidden representations into interpretable biological features without supervision. Across layers and model sizes, SAEs identified over 60 diverse functional annotations encoded in the models activations. This included viral regulatory elements such as the CMV enhancer, despite viral genomes being excluded from training data. Tracing this signal revealed contamination in reference databases, demonstrating that interpretability methods can audit training data and identify hidden data leakage. We then show that Meta-SAEs, trained on the decoder weights of another SAE, can identify conceptual hierarchies encoded in the model, including a more abstract feature related to multiple HIV annotations. We confirmed that the features identified by our SAEs were learned during pretraining through probing a randomly initialised model. Finally, we demonstrate that our SAEs allow us to steer model predictions in biologically meaningful ways, showing that we can use an antibiotic-resistance SAE-feature to steer the model toward the A1408G aminoglycoside-resistance mutation in the ribosomal gene 16S rRNA. Together, these results establish SAEs as a method for both discovery and auditing, providing a toolkit for interpretable and trustworthy genomic foundation models. Readers can explore our findings at https://interpretglm.netlify.app/.

genomics↗

Decoding the physicochemical basis of taxonomy preferences in protein design models

Protein design models have transformed protein engineering by enabling computational exploration of sequence spaces far exceeding experimental capacity. However, their outputs are shaped by both the protein distributions represented in their training corpora and the information available during scoring, so the same model score may reflect backbone-compatible biophysics, taxonomic structure in sequence databases, or other learned regularities rather than protein fitness alone. Here we quantify systematic preferences across 14 protein design models that differ in data modality, training-corpus composition, and scoring context, for comparison grouped as backbone-conditioned, structure plus native-sequence context, or sequence-only. Backbone-conditioned models retain little unexplained taxonomic variance after controlling for protein family and measurable biophysical properties, with residual species variance below 3.3%. In contrast, sequence-only models retain substantial residual taxonomic dependence of 15-20%, indicating that likelihood remains strongly entangled with organism-level sequence statistics. These differences across model classes produce distinct preference landscapes. Backbone-conditioned models organise scores around compactness, packing, and charge, while sequence-only models preserve stronger within-family taxonomic effects. Redesign experiments show that these preferences propagate into generation, shifting templates toward characteristic biophysical profiles rather than uniformly sampling backbone-compatible sequence space. Continued training of ProteinMPNN on ecologically selected extremophile secretomes redirects designed surface chemistry along an acid-base axis while largely preserving structural compatibility and global taxonomic structure. These results show that systematic preferences are not a single failure mode, but separable components arising from scoring context, training-corpus composition, and learned biophysical constraints. Together, they provide a framework for disentangling the sources of model preference and linking them to both scoring behaviour and generated sequence properties.

biochemistry↗