bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.04.09.647986

Fast and Memory-Efficient Dynamic Programming Approach for Large-Scale EHH-Based Selection Scans

Abstract

Haplotype-based statistics are widely used for finding genomic regions under positive selection. At the heart of many such statistics is the computation of extended haplotype homozygosity (EHH), which captures the decay of homozygosity away from a focal site. This computation, repeated for potentially millions of sites, is computationally demanding, as it involves tracking counts of unique haplotypes iteratively over long genomic distances and across many individuals. Because of these computational challenges, existing tools do not scale well when applied to large-scale population datasets, such as the 1000 Genomes Project, or the UK Biobank with 500,000 individuals. Optimizing computation becomes crucial when data sets grow large, especially when handling large sample sizes or generating training data for machine learning algorithms. https://github.com/szpiech/selscan Here, we propose a dynamic programming algorithm that substantially improves runtime and memory usage over existing tools on both real and simulated data. On real phased data, we achieve 5-50x speedup with minimal memory footprint. Our simulations show an even more pronounced performance gap with large populations (up to 15x speedup and 46x memory reduction). EHH-based statistics designed for unphased genotypes run an order of magnitude faster, and multi-parameter support results in 20x runtime improvement. Source code and binaries are available at https://github.com/szpiech/selscan as selscan v2.1.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rahman, A., Smith, T. Q., Szpiech, Z. A.. 2025-04-15. Fast and Memory-Efficient Dynamic Programming Approach for Large-Scale EHH-Based Selection Scans. https://doi.org/10.1101/2025.04.09.647986

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A near-complete chromosome from a bacterial pathogen integrated into the genome of its arthropod vector

Lateral gene transfer (LGT) from organellar to eukaryotic genomes is ubiquitous, resulting in the presence of widespread nuclear mitochondrial DNA segments (NUMTs) and/or plastid gene transfers. While LGT from bacterial associates of eukaryotes is generally less common, many arthropod genomes are littered with LGT originating from Wolbachia (order Rickettsiales) - a vertically transmitted, obligate intracellular symbiont of invertebrates. This contrasts with Rickettsia, a related genus including many important human pathogens, which has not been shown to promulgate LGT within its arthropod vectors. Here, we present the 8.6 Gb genome of a cell line derived from A. variegatum (the tropical bont tick), which is the primary vector of Rickettsia africae (agent of African tick-bite fever). In addition to numerous NUMTs and diverse families of transposable elements, an almost-complete chromosome of R. africae origin was identified in the cell line genome, localised to the putative sex chromosome. Sequencing of field-collected A. variegatum confirmed the presence of this large LGT in wild ticks, although the absence of transferred genes from the R. africae plasmid provided a means to differentiate LGT from genuine rickettsial infections. These findings highlight the imperative to consider the possibility of LGT when screening vectors for pathogens by PCR.

genomics↗

Discovery of two potential new species and two novel bat-coronavirus subgenera (Phyllacovirus and Phyllobecovirus) in the Neotropics

Bats are major natural reservoirs for coronaviruses, yet complete viral genomes from South America remain scarce, limiting evolutionary and taxonomic understanding. Here, we conducted metatranscriptomic sequencing of coronavirus-positive bat samples collected across two ecologically distinct Brazilian biomes: the Atlantic Forest and the semi-arid Caatinga. We recovered seven complete or near-complete genomes belonging to Alphacoronavirus and Betacoronavirus. Phylogenetic and comparative similarity analyses of conserved replicase domains (3CLpro, NiRAN, RdRp, ZBD, HEL1), following International Committee on Taxonomy of Viruses (ICTV) demarcation criteria, revealed significant viral diversity. Within Alphacoronavirus, two genomes from Atlantic Forest phyllostomid bats (Artibeus lituratus and Carollia perspicillata) formed a deeply divergent sister lineage to Amalacovirus, exhibiting a mean amino acid similarity of 76.7% with the reference genome. Within Betacoronavirus, one genome from a Caatinga phyllostomid bat (Artibeus planirostris) clustered within the recently described Ambecovirus clade, displaying 75.9% mean amino acid similarity with mormoopid-associated reference sequences. Based on these divergence levels and non-recombinant genomic architectures, we propose two novel candidate subgenera, Phyllacovirus and Phyllobecovirus, alongside potential novel viral species. Furthermore, our findings demonstrate strong host-associated structuring and biogeographical partitioning of viral lineages across Neotropical biomes. Overall, this study expands the genomic landscape of South American bat coronaviruses and underscores the importance of continuous genomic surveillance at human-wildlife interfaces.

genomics↗

An atlas of eukaryotic centromere architecture reveals recurrent evolutionary dynamics

Centromeres evolved at the root of eukaryotes to segregate chromosomes during cell division. Despite their ancient origin, centromeric DNA sequences evolve rapidly and adopt diverse architectures, including point centromeres, satellite arrays, transposon clusters, and holocentrics. To analyse centromere evolution at a broad scale, we characterised architectures across 325 diverse Darwin Tree of Life genome assemblies. Centromere architecture is evolutionarily labile, and similar configurations arise independently across divergent lineages. In plants and animals, we modelled centromere evolution as a recurrent cycle, in which satellite- and transposon-based architectures interconvert, with independent origins of holocentricity. We curated >23 million satellite repeats comprising 263 families from 165 species. Despite sequence divergence between satellite families, higher order repeats are prevalent, indicating constraint on repeat architecture rather than primary sequence. Satellite arrays are heavily invaded by diverse transposon families, consistent with convergent adaptation to the centromeric niche. In 89 species, transposons themselves constitute the primary centromere structure. We observed centrophilic transposons forming tandem arrays, suggesting mechanisms for satellite regeneration. Our sample includes five independent origins of holocentricity in plants and animals, which vary in association with periodic satellite arrays. We propose that genetic instability, centrophilic transposition, and transmission distortion promote recurrent centromere architectural interconversions during evolution.

genomics↗