bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.06.24.733252

Metrics for Distinguishing Biological and Interventional Change in AI Models

Abstract

Statistical and machine-learning models of longitudinal biological data evaluate change by comparing each new observation against the trajectory implied by prior observations, assuming the process generating that trajectory is stable. We use "data substrate" to mean the underlying structure of the longitudinal data that determines what any such model can recover, independent of its architecture or capacity. When the generating process changes -- whether through a biological transition or through an external intervention -- the prior trajectory ceases to be a valid reference, and extrapolated predictions can be confidently wrong with no internal signal that the reference has failed. A distinct and recognised difficulty is that biological change and interventional change, observed only through serial intertemporal comparison under an assumed trajectory, are readily conflated; existing approaches address this through causal assumptions or hidden-confounder models rather than from the data substrate itself. Here we ask whether the two can be distinguished at the substrate level, and we introduce two subject-level metrics that quantify the geometric signature an interventional change leaves in the data: Curvature Shift, the change in trajectory slope across the event, and Deformation Risk, the departure of post-event observations from the prior-trajectory reference. We evaluate the condition on longitudinal cognitive measurements from 309 human subjects in the Alzheimers Disease Neuroimaging Initiative (ADNI), a large longitudinal dataset containing two distinct, ex-ante-defined regime-change events in the same subjects: a biological transition and an intervention. A model extrapolating the pre-event trajectory assigned the wrong direction of change to roughly two-thirds of post-event observations (post-event sign accuracy 0.341 after the biological event and 0.350 after the intervention, against a chance value of 0.50); only 11% of postbiological-event and 12% of post-intervention readings remained concordant with prior dynamics, and a higher-capacity multilayer perceptron reproduced rather than resolved the error. Curvature Shift was 2.23-fold higher after the biological event (p = 4.4x10-8) and 2.26-fold higher after the intervention (p = 7.4x10-8), and the two metrics were coupled ({rho}= 0.500; 95% CI, 0.407-0.587). Findings replicated on an independent endpoint and survived propensity matching, permutation, 1 and leave-one-out. The metrics detect, per subject, when a fitted models reference has stopped governing the data and whether the departure carries the geometric signature of an interventional change.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ewing, M. A.. 2026-06-29. Metrics for Distinguishing Biological and Interventional Change in AI Models. https://doi.org/10.64898/2026.06.24.733252

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Senescence-associated KRAS upregulation in peripheral T cells links to premature coronary artery disease

Aims: Premature coronary artery disease (PCAD) lacks specific molecular drivers, and the role of immunosenescence is unclear. We investigated whether aging-related gene dysregulation in T cells contributes to PCAD. Methods: We combined bulk transcriptomics of PBMCs from 12 PCAD patients and 21 controls, single-cell RNA sequencing of PBMCs and human atherosclerotic plaques, weighted gene co-expression network analysis, gene perturbation network analysis, and molecular docking. Results: KRAS was identified as a hub gene intersecting PCAD-associated genes and aging-related genes. Single-cell analysis showed KRAS upregulation predominantly in effector CD8+ T cells, which exhibited the highest senescence scores that were further elevated in disease. Network perturbation of KRAS strongly impacted the cell killing pathway. KRAS-high effector CD8+ T cells were detected in coronary and carotid plaques, displaying enhanced cytotoxicity, exhaustion, and senescence features. Additionally, a candidate small molecule was computationally predicted to bind inactive KRAS. Conclusions: Elevated KRAS expression in senescent, cytotoxic CD8+ T cells is associated with PCAD, bridging immunosenescence and premature atherosclerosis. This finding provides a novel biomarker candidate and potential therapeutic entry point, awaiting further functional validation.

bioinformatics↗

Targeted finetuning enables co-folding models to learn ligand-induced protein conformational states

Advances in protein structure prediction have enabled all-atom protein-ligand co-folding models that predict bound conformations directly from sequence and small-molecule structure. However, these models often fail to generalize to novel binding sites or alternative protein conformational states, limiting their utility for chemical biology and drug discovery. Here we show this limitation reflects training data bias rather than architectural constraints and can be overcome through targeted finetuning. Using ten previously unseen X-ray structures of Werner (WRN) helicase from a drug discovery program, we finetune Boltz-1 to learn both an allosteric binding site and a large conformational change locking the enzyme in an inactive state, while preserving accuracy on the ATP-bound state. The finetuned model generalizes to different chemical series and transfers the conformational logic across RecQ-family helicases in a binding-site sequence-dependent manner. This approach provides a blueprint for adapting foundation models as new structural and mechanistic data emerge, enabling co-folding networks to capture ligand-induced conformational switches and binding poses absent from their training data but central to biological regulation and therapeutic intervention.

bioinformatics↗

Benchmarking single-cell foundation models for aging biology

Single cell foundation models (scFMs) provide representations of cellular states, but their utility across biological questions in aging research remains unclear. We established a benchmark of cellular representations for aging research, evaluating ten general-purpose scFMs, three aging-specific models and conventional methods across five biological questions using more than 2.5 million single cell transcriptomes. Using frozen pretrained representations, Geneformer performed best among scFMs for chronological age prediction and age pseudotime concordance, although 2,000 highly variable genes achieved higher mean performance. Several scFMs captured positive molecular age shifts across three disease contexts, consistent with reported aging-associated changes. SCimilarity performed well for rare cellular state identification across out-of-distribution datasets, exceeding aging specific models and conventional baselines. At the gene level, scGPT showed the highest recovery of reference TF target interactions, including aging-related regulatory hubs. Overall, scFMs supported diverse aging analyses, but performance depended on the biological question, highlighting their utility for rare cellular state identification and regulatory analysis.

bioinformatics↗