bioRxiv Science⌕ Search

Biology subjects

Calvanese, F.

Publications and source records attributed to Calvanese, F..

4 recordsLinked to original sources

Steering Sequence Generation in Protein Language Models through Iterative Lookback Monte Carlo Sampling

Protein language models (pLMs) leverage large-scale evolutionary data to generate novel sequences, but steering generation toward desired physicochemical properties without sacrificing diversity remains a major challenge. Existing approaches often induce severe diversity loss or require computationally expensive retraining. We introduce Iterative Lookback Monte Carlo (ILMC), a training-free inference-time sampling strategy that interleaves autoregressive elongation with Metropolis-Hastings refinement to approximate sampling from a maximum-entropy target distribution balancing generative quality and steering objectives. We show theoretically that this target distribution is entropy-maximizing under fixed generative quality and steering constraints, and empirically that ILMC produces more diverse samples than standard autoregressive baselines at matched generative quality. Using simple steering potentials, ILMC improves desired molecular properties, including generating proteins with up to 12{degrees}C higher predicted melting temperature than compute-matched alternative strategies. ILMC naturally applies to classifier-guided steering, where it outperforms purely autoregressive guidance in diversity while maintaining comparable enrichment of target properties. We validate ILMC on family-specific pLMs and on the multi-family model ProGen3.

bioinformatics↗

Integrating experimental feedback improves generative models for biological sequences

Generative probabilistic models have shown promise in designing artificial RNA and protein sequences but often suffer from high rates of false positives, where sequences predicted as functional fail experimental validation. To address this critical limitation, we explore the impact of reintegrating experimental feedback into the model design process. We propose a likelihood-based reintegration scheme, which we test through extensive computational experiments on both RNA and protein datasets, as well as through wet-lab experiments on the self-splicing ribozyme from the group I intron RNA family where our approach demonstrates particular efficacy. We show that integrating recent experimental data enhances the models capacity of generating functional sequences (e.g. from 6.7% to 63.7% of active designs at 45 mutations). This feedback-driven approach thus provides a significant improvement in the design of biomolecular sequences by directly tackling the false-positive challenge.

bioinformatics↗

Expanding the space of self-reproducing RNAs using probabilistic generative models

Estimating the plausibility of RNA self-reproduction is central to origin-of-life scenarios but self-reproduction has been shown in only a handful of systems. Here, we populated a vast sequence space of ribozymes using statistical covariation models and secondary structure prediction. Experimentally assayed sequences were found active as far as 65 mutations from a reference natural sequence. The number of potentially generated sequences together with the experimental success rate indicate that at least [~]1039 such ribozymes may exist. Randomly sampled artificial ribozymes exhibited autocatalytic self-reproduction akin to the reference sequence. The combination of high-throughput screening and probabilistic modeling considerably improves our estimation of the number of self-reproducing systems, paving the way for a statistical approach to the origin of life.

biochemistry↗

Towards Parsimonious Generative Modeling of RNA Families

Generative probabilistic models emerge as a new paradigm in data-driven, evolution-informed design of biomolecular sequences. This paper introduces a novel approach, called Edge Activation Direct Coupling Analysis (eaDCA), tailored to the characteristics of RNA sequences, with a strong emphasis on simplicity, efficiency, and interpretability. eaDCA explicitly constructs sparse coevolutionary models for RNA families, achieving performance levels comparable to more complex methods while utilizing a significantly lower number of parameters. Our approach demonstrates efficiency in generating artificial RNA sequences that closely resemble their natural counterparts in both statistical analyses and SHAPE-MaP experiments, and in predicting the effect of mutations. Notably, eaDCA provides a unique feature: estimating the number of potential functional sequences within a given RNA family. For example, in the case of cyclic di-AMP riboswitches (RF00379), our analysis suggests the existence of approximately 1039 functional nucleotide sequences. While huge compared to the known < 4, 000 natural sequences, this number represents only a tiny fraction of the vast pool of nearly 1082 possible nucleotide sequences of the same length (136 nucleotides). These results underscore the promise of sparse and interpretable generative models, such as eaDCA, in enhancing our understanding of the expansive RNA sequence space.

bioinformatics↗