bioRxiv Science⌕ Search

Biology subjects

Selberg, A.

Publications and source records attributed to Selberg, A..

3 recordsLinked to original sources

HyphAeon: Attention on Evolution Across Deep Time Transforms Comparative Genomics

Detecting Darwinian natural selection is fundamental to evolutionary biology and functional genomics, yet standard methods based on phylogenetic models that estimate the ratio of non-synonymous to synonymous substitution rates (dN/dS) fail to scale with modern genomic volumes. Fitting continuous-time Markov substitution matrices across dense trees with hundreds of species requires extensive compute, forcing comparative genomics to rely on aggressive taxon subsampling or static whole-tree summaries that dilute transient adaptive bursts. Here we present HyphAeon, a lightweight (~1.91M parameter backbone, 2.46M across the full multi-task suite) phylogeny-informed foundation transformer trained to amortize the detection of episodic diversifying selection across 742-species mammalian coding alignments (17,186 genes, 9.77 x 10^6 codons). HyphAeon approaches the discriminative accuracy of numerical maximum-likelihood selection tests (MEME) across episodic burst regimes (ROC-AUC up to 0.942, mean 0.659; empirical Precision-Recall lift up to 25.8x, mean 7.7x; rank concordance up to rho = 0.983) while executing >1,000x faster per locus (averaging <1 ms per site across genome-wide scans) and >10,000x faster at proteome scale, generalizing outside its mammalian training distribution without retraining. Beyond accelerating classical tests, embedding molecular evolution into a differentiable geometric latent space enables analytical capabilities inaccessible to static dN/dS models: (1) targeted alignment artifact correction via counterfactual attribution; (2) macromolecular contact recovery and multi-site epistatic sectors (CESI); (3) directional phenotype-to-genotype attribution in lineage space (PARS); and (4) continuous temporal surveillance regression that tracks positive sweep velocities across longitudinal cohorts (evaluated across 12,167 timestamped genomes and benchmarked against external frequencies from >9.34 million genomes), rescuing adaptive substitutions obscured by post-fixation dilution. By bridging statistical phylogenetics with geometric representation learning, HyphAeon establishes comparative genomics as an interactive, high-throughput computational framework for evolutionary discovery.

bioinformatics↗

BUSTED-PH: Isolating the genomic signatures of convergent phenotypes

Convergent evolution offers a natural test for adaptive predictability, yet pinpointing its molecular basis remains difficult. Current methods often grapple with distinguishing adaptive convergence from background noise, either by demanding overly stringent identical substitutions or relying on low-resolution evolutionary rate shifts. Here we introduce BUSTED-PH (Branch-site Unrestricted Statistical Test for Episodic Diversification - Phenotype), a branch-site codon model that detects phenotype-associated episodic diversifying selection. Unlike standard approaches, BUSTED-PH explicitly contrasts selective regimes between phenotype-positive (foreground) and phenotype-negative (background) lineages, effectively winnowing out spurious associations driven by pervasive background adaptation. BUSTED-PH has already been applied in numerous independent studies, and here we rigorously validate it using canonical positive controls and simulations to confirm high power and strict false positive control. A genome-wide scan of 120 mammalian species using strict statistical criteria (FDR [&le;] 0.01) identified 72 genes associated with echolocation; while recovering paradigmatic auditory drivers (e.g., Prestin, TMC1 ), we also uncover novel candidates in physiological support systems ranging from lipid homeostasis to neural development. A parallel analysis of mammalian gigantism identifies 91 genes linked to musculoskeletal reinforcement, organ size governance, and genomic integrity, characterizing the molecular adaptations required to support massive body size. As projects like Zoonomia and the Vertebrate Genomes Project expand comparative datasets, BUSTED-PH provides a robust framework for dissecting the genetic architecture of complex convergent traits.

evolutionary biology↗

Minus the Error: Estimating dN/dS and Testing for Natural Selection in the Presence of Residual Alignment Errors

Positive selection is an evolutionary process which increases the frequency of advantageous mutations because they confer a fitness benefit. Inferring the past action of positive selection on protein-coding sequences is fundamental for deciphering phenotypic diversity and the emergence of novel traits. With the advent of genome-wide comparative genomic datasets, researchers can analyze selection not only at the level of individual genes but also globally, delivering systems-level insights into evolutionary dynamics. However, genome-scale datasets are generated with automated pipelines and imperfect curation that does not eliminate all sequencing, annotation, and alignment errors. Positive selection inference methods are highly sensitive to such errors. We present BUSTED-E: a method designed to detect positive selection for amino acid diversification while concurrently identifying some alignment errors. This method builds on the flexible branch-site random effects model (BUSTED) for fitting distributions of dN/dS, with a critical modification: it incorporates an "error-sink" component to represent an abiological evolutionary regime. Using several genome-scale biological datasets that were extensively filtered using state-of-the art automated alignment tools, we show that BUSTED-E identifies pervasive residual alignment errors, produces more realistic estimates of positive selection, reduces bias, and improves biological interpretation. The BUSTED-E model promises to be a more stringent filter to identify positive selection in genome-wide contexts, thus enabling further characterization and validation of the most biologically relevant cases.

evolutionary biology↗