bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.05.21.726777

Fast pairwise coalescence enables gene-resolution scans for recent selection in diverse human populations

Abstract

Identifying the genetic changes that shaped recent human adaptation depends on our ability to detect selection from genomic data. Summary statistics from haplotype scans have been widely used for that purpose, aggregating genetic signal over windows, though resolution is limited by linkage and their power may diminish as sweeps approach fixation, as in the case of the integrated haplotype score (iHS). Ancient DNA based scans recover signal by analysing time-series trajectories, but the majority of human populations fall outside the geographic range of any existing ancient DNA dataset. Pairwise coalescence times provide a way to complement statistics and can be applied to any modern cohort, yet computing them densely enough at cohort scale poses a computational challenge due to the quadratic growth in the number of haplotype pairs. We introduce gamma_smc_cu, a GPU implementation of the Gamma-SMC algorithm (Schweiger and Durbin, 2023) for pairwise time-to-the-most-recent-common-ancestor (TMRCA) inference. Applied to the 1000 Genomes Project (3,202 phased samples, corresponding to 6,404 haplotypes; 829,638 within-population pairs across 26 populations and five different continental ancestries; [~]1012 per-site posterior evaluations), it yields a gene-level TMRCA landscape of 17,823 autosomal protein-coding genes after masking for segmental duplications. The scan recovers well-known sweeps (LCT, SLC24A5, EDAR, FADS1, HERC2, ABCC11) and, combined with a depleted-to-enriched variant-class profile, resolves haplotype-block signals down to the gene level. Of seven case studies, two are developed in the main text -- GRK2 /ADRBK1 (chr11q13.2; SAS+EUR) and TREML1 /TREM2 (chr6p21.1) -- and the remaining five (IFIH1 chr2q24/IBS, CCDC92 chr12q24/CDX, SLC6A15 chr12q21/CHS, BPIFA2 chr20q11/GIH, CLEC6A chr12p13/CDX) are presented in the Supplementary Information (SI). Notably, TREML1 /TREM2 is a shared out-of-Africa signal -- ranked below the within-population 1% tail in 16 of 19 non-African 1000 Genomes panels that PopHumanScan and five landmark haplotype-based scans miss. A previous 10 kb-windowed-mean iHS scan dilutes the cluster of extreme sites packed inside the [~]5 kb gene bodies, while our own gene-level iHS independently recovers the locus in three South Asian panels (BEB, STU, ITU; top 0.4% genome-wide). We cross-validate the seven cases against the 9.7 million per-variant selection posteriors from a recent West-Eurasian ancient DNA scan. BPIFA2 is detected concordantly (s {approx} 1.8% per generation). GRK2 and CCDC92 reach detection threshold in flanking variants but not within their own gene bodies, while the TREML1 /TREM2 cluster falls below it. To calibrate novelty, we review the candidate landscape against an expanded eight-catalog set spanning curated haplotype scans, the largest current West-Eurasian ancient-DNA leads, and a recent 26-population iHS refinement; the vast majority of our loci overlap at least one prior entry, and only a handful -- including TREML1 /TREM2 -- remain unflagged. The contributions of this work are gene-level resolution, systematic ancient DNA cross-validation, and a reusable TMRCA landscape that complements aDNA panels.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Korfmann, K., Mathieson, S.. 2026-05-22. Fast pairwise coalescence enables gene-resolution scans for recent selection in diverse human populations. https://doi.org/10.64898/2026.05.21.726777

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

RELAX does not reproduce its own estimates at default settings, and its output does not show it

Selection-intensity estimates from RELAX are reported as a point value of K with a likelihood-ratio P. We report that, at default settings and on data of ordinary size, the program does not reproduce its own fits. Of 27 enzyme entries refitted under two optimiser configurations, none reproduced its log-likelihood to within 0.01 units; the median change was 103 units, the largest over 3,400, and four verdicts reversed. Eighty null orthologues reproduced none. A byte-identical command returned a distinct likelihood on every repetition, single-threaded, across three releases, and on alignments simulated under the fitted model, where 3.3 per cent of replicates reproduced. The documented random-number seed never reaches the generator when assigned on the command line, yet reads back as the value supplied. PAML localises the cause: its two-ratio model, without site classes, reproduced its log-likelihood for all 288 genes; its site-class models agreed for 27 to 67 per cent. The instability follows the mixture over sites, not the program. The output does not show it: 46 of 410 fits ended with a negative likelihood-ratio statistic, impossible under convergence, and 123 of 410 report a K re-estimated under a domain restriction rather than the unconstrained maximum. Of 234 published studies using RELAX, none reported a seed. Seeding while holding the thread count at one reproduced sixty of sixty runs on twenty genes under two releases; the seed alone reproduced none of five, and no documentation states the second condition. We recommend that fits be repeated and their dispersion published.

evolutionary biology↗

Sequential accumulation of adaptive alleles forms an inversion supergene in deer mice

Supergenes are clusters of co-inherited loci that affect multiple or complex phenotypes. Despite the growing number of chromosomal inversions identified as supergenes in natural populations, their molecular basis and evolutionary history often remain obscure. Here, we identified two candidate genes, Slc45a2 and Npr3, within a 41-Mb inversion supergene in the deer mouse (Peromyscus maniculatus) that respectively drive darker coats and longer tails - two traits associated with forest adaptation. Mice homozygous for the inversion (inv/inv) exhibit elevated Slc45a2 expression in melanocytes relative to the congenic standard genotype (std/std), disrupting pheomelanin production. In parallel, downregulation of Npr3 in inv/inv mouse growth plates prolongs postnatal growth of caudal vertebrae, resulting in tail elongation. Population-level analyses further implicate that this supergene arose through the subsequent accumulation of the Npr3 allele within the inversion, rather than by capturing all beneficial mutations at its origin.

evolutionary biology↗

Toxin structure shapes palatability in a chemically defended butterfly

The toxicity of chemical defences is well studied, but the potential contribution of compound structure to predator deterrence remains largely unexplored. Whether predation acts more strongly on toxicity or unpalatability remains largely untested, partly because few systems allow toxin structure to vary independently of quantity. Heliconius sara larvae provide such a system: those reared on Passiflora auriculata sequester cyclopentenyl cyanogenic glucosides (CGs), while those reared on P. biflora biosynthesise comparable quantities of aliphatic CGs. Using two invertebrate predators, Camponotus floridanus ants and Hierodula membranacea mantids, we tested whether this structural difference affects palatability independent of toxicity. Mantids rejected larvae with cyclopentenyl CGs more often than larvae with aliphatic CGs, despite no detectable difference in total CG content. This pattern was mirrored in extract-based assays with ants, independently of cyanide release: extracts with cyclopentenyl CGs remained deterrent, while extracts with aliphatic CGs did not differ in deterrence from water. Live larvae, by contrast, elicited similar responses from ants regardless of CG structure. These results show that variation in toxin structure can strongly affect palatability, with some compounds conferring greater protection than others. This demonstrates the importance of chemical structural diversity in the evolution of chemical defences.

evolutionary biology↗