bioRxiv ScienceSearch

Biology subjects

Zamponi, F.

Publications and source records attributed to Zamponi, F..

2 recordsLinked to original sources

Efficient generative modeling of protein sequences using simple autoregressive models

Generative models emerge as promising candidates for novel sequence-data driven approaches to protein design, and for the extraction of structural and functional information about proteins deeply hidden in rapidly growing sequence databases. Here we propose simple autoregressive models as highly accurate but computationally extremely efficient generative sequence models. We show that they perform similarly to existing approaches based on Boltzmann machines or deep generative models, but at a substantially lower computational cost. Furthermore, the simple structure of our models has distinctive mathematical advantages, which translate into an improved applicability in sequence generation and evaluation. Using these models, we can easily estimate both the model probability of a given sequence, and the size of the functional sequence space related to a specific protein family. In the case of response regulators, we find a huge number of ca. 1068 sequences, which nevertheless constitute only the astronomically small fraction 10-80 of all amino-acid sequences of the same length. These findings illustrate the potential and the difficulty in exploring sequence space via generative sequence models.

bioinformatics

Aligning biological sequences by exploiting residue conservation and coevolution

Aligning biological sequences belongs to the most important problems in computational sequence analysis; it allows for detecting evolutionary relationships between sequences and for predicting biomolecular structure and function. Typically this is addressed through profile models, which capture position-specificities like conservation in sequences, but assume an independent evolution of different positions. RNA sequences are an exception where the coevolution of paired bases in the secondary structure is taken into account. Over the last years, it has been well established that coevolution is essential also in proteins for maintaining three-dimensional structure and function; modeling approaches based on inverse statistical physics can catch the coevolution signal and are now widely used in predicting protein structure, protein-protein interactions, and mutational landscapes. Here, we present DCAlign, an efficient approach based on an approximate message-passing strategy, which is able to overcome the limitations of profile models, to include general second-order interactions among positions and to be therefore universally applicable to protein- and RNA-sequence alignment. The potential of our algorithm is carefully explored using well-controlled simulated data, as well as real protein and RNA sequences.

bioinformatics