bioRxiv Science⌕ Search

Biology subjects

Karim, L. M.

Publications and source records attributed to Karim, L. M..

2 recordsLinked to original sources

Real-time accessible phylogenetics for every highly sampled virus

The scale of viral genome sequencing has outpaced the phylogenetic tools traditionally used to analyze it, as highlighted by the COVID-19 pandemic. We present viral_usher, a unified framework for scalable viral phylogenetics built on UShER. viral_usher is a containerized command-line tool that constructs mutation-annotated trees directly from public sequence repositories with minimal user input, building phylogenies of tens of thousands of genomes in minutes. Applying it across the International Nucleotide Sequence Database Collaboration, we assembled viral_usher_trees, a repository of 446 phylogenies spanning 163 well-sequenced viral species, rebuilt automatically monthly as new genomes are deposited. We extended Taxonium from a tree viewer into a web platform supporting in-browser phylogenetic placement and de novo tree construction, so that users can upload sequences and contextualize them within global phylogenies without local computational infrastructure. Because every tree is built by the same procedure, the repository enables comparative analyses across the breadth of viral diversity. We demonstrate the utility of this resource by asking what factors shape viral mutation spectra. We found that replication machinery, captured as Baltimore class, explains 46% of the variance across 162 viral genomes, while host taxon and envelope status together explain under 5%. These resources provide an extensible platform for real-time genomic epidemiology and for comparative evolutionary analysis across viral pathogens. Resources and code are freely available at https://taxonium.org/, https://github.com/lilymaryam/spectrum_analysis, https://github.com/AngieHinrichs/viral_usher_trees, and https://github.com/AngieHinrichs/viral_usher.

bioinformatics↗

Panmap: Scalable phylogeny-guided alignment, genotyping, and placement on pangenomes

Pangenomes capture population-level variation but remain computationally challenging at scale. We present Panmap, a tool that leverages evolutionary structure to place, align, and genotype sequencing reads against mutation-annotated pangenomes containing up to millions of genomes. Panmap introduces a phylogenetically compressed k-mer index that stores only sequence differences along branches, enabling efficient comparison of reads to both sampled genomes and inferred ancestors. This approach reduces index size by up to 600-fold and construction time by over three orders of magnitude relative to existing tools. Panmap places a 100x coverage SARS-CoV-2 sample onto 20,000 genomes in 0.4 seconds and onto 8 million genomes in under two minutes. Furthermore, it enables accurate haplotype identification and abundance estimation in metagenomic samples and sensitive placement of ancient environmental DNA without prior alignment. Our approach makes large-scale pangenomes directly amenable to read mapping, genome assembly, alignment-free phylogenetic placement, and metagenomic analysis.

bioinformatics↗