bioRxiv · 10.1101/2024.01.27.577580
nail: software for high-speed, high-sensitivity protein sequence annotation
Abstract
Fast is fine, but accuracy is final. --Wyatt Earp BackgroundProfile hidden Markov models (pHMMs) deliver state-of-the-art sensitivity for sequence annotation, but the Forward/Backward algorithm fills a dynamic programming matrix sized by the product of model and sequence lengths, making pHMM search slower than fast heuristic alignment tools like MMseqs2 by an order of magnitude or more. ResultsWe introduce nail, which approximates Forward/Backward by computing only a sparse cloud of high-probability matrix cells, recovering accurate pHMM scores, E-values, and alignments at a fraction of the cost. nail annotates the ~2.4 billion protein MGnify metagenomic dataset with all of Pfam in 73.9 hours on a single 48 core machine, recovering nearly all of HMMERs recall advantage over MMseqs2, with run time ~8.7x faster than HMMER3. Detailed analysis of HMMER-only multi-domain hits suggests that many of these matches missed by nail are the result of accumulating score across short, often fragmentary alignments to repetitive regions of the target, consistent with spurious hits rather than genuine homology. We also derive a closed-form approximation for single-sequence E-value calibration, eliminating a per-model simulation step. nail is released under the open BSD-3-clause license at https://github.com/TravisWheelerLab/nail.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Roddy, J. W., Rich, D. H., Wheeler, T. J.. 2024-01-30. nail: software for high-speed, high-sensitivity protein sequence annotation. https://doi.org/10.1101/2024.01.27.577580
Cite the original work for its findings. Save a collection to share your selection of sources.