bioRxiv · 10.1101/2024.05.24.595840
Sequence alignment using large protein structure alphabets doubles sensitivity to remote homologs
Abstract
Recent breakthroughs in protein fold prediction from amino acid sequences have unleashed a deluge of new structures, raising new opportunities for expanding insights into the universe of proteins and pursuing practical applications in bio-engineering and therapeutics while also presenting new challenges to protein search and analysis algorithms. Here, I describe Reseek, a protein alignment algorithm which improves sensitivity in protein homolog detection compared to state-of-the-art methods including DALI, TM-align and Foldseek, with improved speed over Foldseek, the fastest previous method. Reseek is based on alignment of sequences where each residue in the protein backbone is represented by a letter in a novel "mega-alphabet" of 85,899,345,920 ([~] 1011) distinct states. Code is available at https://github.com/rcedgar/reseek.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Edgar, R. C.. 2024-05-27. Sequence alignment using large protein structure alphabets doubles sensitivity to remote homologs. https://doi.org/10.1101/2024.05.24.595840
Cite the original work for its findings. Save a collection to share your selection of sources.