bioRxiv · 10.1101/2020.05.18.101295
Aligning biological sequences by exploiting residue conservation and coevolution
Abstract
Aligning biological sequences belongs to the most important problems in computational sequence analysis; it allows for detecting evolutionary relationships between sequences and for predicting biomolecular structure and function. Typically this is addressed through profile models, which capture position-specificities like conservation in sequences, but assume an independent evolution of different positions. RNA sequences are an exception where the coevolution of paired bases in the secondary structure is taken into account. Over the last years, it has been well established that coevolution is essential also in proteins for maintaining three-dimensional structure and function; modeling approaches based on inverse statistical physics can catch the coevolution signal and are now widely used in predicting protein structure, protein-protein interactions, and mutational landscapes. Here, we present DCAlign, an efficient approach based on an approximate message-passing strategy, which is able to overcome the limitations of profile models, to include general second-order interactions among positions and to be therefore universally applicable to protein- and RNA-sequence alignment. The potential of our algorithm is carefully explored using well-controlled simulated data, as well as real protein and RNA sequences.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Muntoni, A. P., Pagnani, A., Weigt, M., Zamponi, F.. 2020-05-20. Aligning biological sequences by exploiting residue conservation and coevolution. https://doi.org/10.1101/2020.05.18.101295
Cite the original work for its findings. Save a collection to share your selection of sources.