bioRxiv Science⌕ Search

Biology subjects

Callan, D.

Publications and source records attributed to Callan, D..

6 recordsLinked to original sources

PyEuk: a tool suite for catalogue-free multilocus typing

Multilocus sequence typing anchors molecular epidemiology, but traditional frameworks require centrally curated allele catalogues. For emerging and uncultivable eukaryotic parasites, maintaining these databases is impractical, leaving surveillance reliant on fragmented, assay-specific scripts. PyEuk eliminates this bottleneck by providing an open, catalogue-free suite that calls microhaplotypes directly from sequence differences relative to a reference within data-defined genomic windows. When amplicon coordinates are uncharacterized or unpublished, PyEuk reconstructs target panels de novo from raw read coverage peaks mapped to a draft assembly. Across benchmark cohorts spanning Cyclospora cayetanensis and Plasmodium vivax, PyEuk recovers epidemiological structure established by tracebacks, geography, and clinical recurrence without organism-specific tuning. In foodborne outbreaks, it resolves independent transmission chains using either curated or de novo panels and scales to national surveillance archives exceeding 8,000 isolates. In P. vivax malaria, its weighted identity-by-state distance separates continental lineages, discriminates liver-stage relapses from reinfections, and delineates transmission clusters. Rather than forcing an arbitrary partition on continuous variation, PyEuk evaluates bootstrap stability, reporting supported cluster count ranges alongside reproducible transmission cores. PyEuk provides a portable, reproducible foundation for eukaryotic pathogen surveillance.

genomics↗

HyphAeon: Attention on Evolution Across Deep Time Transforms Comparative Genomics

Detecting Darwinian natural selection is fundamental to evolutionary biology and functional genomics, yet standard methods based on phylogenetic models that estimate the ratio of non-synonymous to synonymous substitution rates (dN/dS) fail to scale with modern genomic volumes. Fitting continuous-time Markov substitution matrices across dense trees with hundreds of species requires extensive compute, forcing comparative genomics to rely on aggressive taxon subsampling or static whole-tree summaries that dilute transient adaptive bursts. Here we present HyphAeon, a lightweight (~1.91M parameter backbone, 2.46M across the full multi-task suite) phylogeny-informed foundation transformer trained to amortize the detection of episodic diversifying selection across 742-species mammalian coding alignments (17,186 genes, 9.77 x 10^6 codons). HyphAeon approaches the discriminative accuracy of numerical maximum-likelihood selection tests (MEME) across episodic burst regimes (ROC-AUC up to 0.942, mean 0.659; empirical Precision-Recall lift up to 25.8x, mean 7.7x; rank concordance up to rho = 0.983) while executing >1,000x faster per locus (averaging <1 ms per site across genome-wide scans) and >10,000x faster at proteome scale, generalizing outside its mammalian training distribution without retraining. Beyond accelerating classical tests, embedding molecular evolution into a differentiable geometric latent space enables analytical capabilities inaccessible to static dN/dS models: (1) targeted alignment artifact correction via counterfactual attribution; (2) macromolecular contact recovery and multi-site epistatic sectors (CESI); (3) directional phenotype-to-genotype attribution in lineage space (PARS); and (4) continuous temporal surveillance regression that tracks positive sweep velocities across longitudinal cohorts (evaluated across 12,167 timestamped genomes and benchmarked against external frequencies from >9.34 million genomes), rescuing adaptive substitutions obscured by post-fixation dilution. By bridging statistical phylogenetics with geometric representation learning, HyphAeon establishes comparative genomics as an interactive, high-throughput computational framework for evolutionary discovery.

bioinformatics↗

Beyond Invariable Sites: Using Evolutionary Stasis to Map Multi-Layered Constraints on the Evolution of Viral and Mammalian Genomes

The quantification of genomic conservation has progressed from foundational statistical modeling of evolutionary rates to state-of-the-art phylogeny-aware deep learning architectures. Yet, a fundamental resolution gap remains whenever evolutionary rates closely approach the "zero-rate origin," where standard selection inference tools will essentially ignore signals of extreme purifying section at invariant genome sites. We present B-STILL (Bayesian Significance Test of Invariant Low Likelihoods), a hierarchical Bayesian framework designed to resolve the selective landscape of protein-coding data by leveraging gene-level calibration and codon-site specific evolutionary opportunity. This framework is based on computationally efficient approximations using codon-substitution models which are scalable to alignments with thousands of sequences. By explicitly tuning the stasis radius around the near-zero evolutionary-rate regime, B-STILL distinguishes between stochastic invariance and functional constraint, identifying Evolutionary Stasis Anchors (ESAs) where the upper bound on permitted evolutionary change is statistically anomalous relative to the background of the gene. This hierarchical approach provides a signature of functional or structural constraint that is often difficult to detect using other tools. Validation against extensive pathogen and clinical databases confirms that ESAs are predictors of biological fitness and disease potential. Collectively, we identified thousands of significantly clustered ESAs that precisely footprint both known functional domains and currently uncharacterized structural motifs in mammalian and viral genomes. These findings establish B-STILL as a scalable statistical framework for high-resolution genomic annotation, transforming formerly ignored invariant genome and protein sites into informative markers of extreme purifying selection across both well-characterized and uncharacterized protein-coding genes from different domains of life.

evolutionary biology↗

CAPHEINE, or everything and the kitchen sink: a workflow for automating selection analyses using HyPhy

SummaryHere we present CAPHEINE, a computational workflow that starts with a set of unaligned pathogen sequences and a reference genome and performs a comprehensive exploratory evolutionary analysis of the input data. CAPHEINE pairs nicely with studies of site-level selection dynamics, gene-level positive selection, and lineage-specific shifts in selective pressure. Our workflow is portable across Mac OS, Windows, and Linux, allowing researchers to focus on results. Availability and ImplementationCAPHEINE is freely available at https://github.com/veg/CAPHEINE, along with a set of usage instructions.

bioinformatics↗

Standardizing RNA-seq Analysis of Fungal Pathogens Using BRC-Analytics and Agentic AI: A Candidozyma auris Case Study

Candidozyma auris has emerged as a critical global health threat due to multidrug resistance and healthcare-associated transmission. While RNA-seq has become the primary tool for studying C. auris pathogenesis, inconsistent use of reference genomes and bioinformatics tools complicate cross-study comparisons. Here we demonstrate how BRC-Analytics, a platform for pathogen genomics, combined with an agentic AI assistant, enables reproducible RNA-seq analysis. By re-analyzing data from two publications we achieved near-perfect correlation with published results despite annotation version differences. We addressed provenance challenges associated with using AI agents with Galaxy by forcing them to invoke Galaxys native tools rather than manipulating data directly. For custom analyses outside Galaxys toolset, we provide standalone JupyterLite notebooks that reproduce our analysis without AI involvement. This framework--combining AI-assisted automation with rigorous provenance tracking--establishes a template for standardized, reproducible fungal pathogen genomics. To the best of our knowledge, this is the first example of integration between public data repositories, reproducible analysis workflows, and agentic AI tools. Our subsequent efforts will focus on improving the seamlessness of this integration.

genomics↗

From data to publication in a browser with BRC-Analytics: Evolutionary dynamics of coding overlaps in measles virus

The analytical landscape of pathogen research is often fragmented, hindering transparency and reproducibility due to diverse genomic data sources, numerous software tools, and suboptimal integration methods. Here we introduce BRC-analytics, a novel browser-based environment that unifies authoritative sources of genomic data with community-curated best analysis practices on a freely accessible public computational infrastructure. We demonstrate its capabilities by analyzing the evolutionary dynamics within the P/V/C locus of the measles virus, a complex system involving overlapping coding regions and RNA editing. Our analysis, conducted entirely within BRC-analytics, reveals asymmetric evolution of the locuss reading frames under distinct selective pressures. BRC-analytics streamlines the entire research process--from data collection and primary analysis (e.g., variant calling) to interpretation (e.g., using integrated JupyterLite notebooks and LLMs) and publication--into a single web browser session. This eliminates the need for local installations and manual data transfers, implicitly tracking provenance and ensuring reproducibility. The platforms goal is to provide true data-to-publication functionality, making advanced pathogen genomics accessible to a broader research community regardless of their computational expertise or infrastructure access.

genomics↗