bioRxiv · 10.1101/2025.03.26.645543
Columba: Fast Approximate Pattern Matching withOptimized Search Schemes
Abstract
Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first is based on the bidirectional FM-index. The second, Columba RLC, employs run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Through extensive benchmarking, Columba outperforms existing lossless aligners in speed, particularly at higher error rates. Tests on the human genome and bacterial and human pan-genome datasets demonstrate Columbas robustness and efficiency. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Renders, L., Depuydt, L., Gagie, T., Fostier, J.. 2025-03-31. Columba: Fast Approximate Pattern Matching withOptimized Search Schemes. https://doi.org/10.1101/2025.03.26.645543
Cite the original work for its findings. Save a collection to share your selection of sources.