bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.07.28.667162

ACMTF-R: supervised multi-omics data integration uncovering shared and distinct outcome-associated variation

Abstract

The rapid growth of high-dimensional biological data has necessitated advanced data fusion techniques to integrate and interpret complex multi-omics and longitudinal datasets. Shared and unshared structure across such datasets can be identified in an unsupervised manner with Advanced Coupled Matrix and Tensor Factorization (ACMTF), but this cannot be related to an outcome. Conversely, N-way Partial Least Squares (NPLS) is supervised and captures outcome-associated variation but cannot identify shared and unshared structure. To bridge the gap between data exploration and prediction, we introduce ACMTF-Regression (ACMTF-R), an extension of ACMTF that incorporates a regression step, allowing for the simultaneous decomposition of multi-way data while explicitly capturing variation associated with a dependent variable. We present a detailed mathematical formulation of ACMTF-R, including its optimisation algorithm and implementation. Through extensive simulations, we systematically evaluate its ability to recover a small y-related component shared between multiple blocks, its robustness to noise, and the impact of the tuning parameter ({pi}) which controls the balance between data exploration and outcome prediction. Our results demonstrate that ACMTF-R can robustly identify the y-related component, correctly identifying outcome-associated shared and distinct variation, distinguishing it from existing approaches such as NPLS and ACMTF. The development of ACMTF-R was motivated by a real-world dataset investigating how maternal pre-pregnancy BMI affects the human milk microbiome, human milk metabolome, and infant faecal microbiome. Emerging evidence suggests that inter-generational transfer of maternal obesity may affect multiple omics layers, highlighting the need to identify outcome-associated variation. The applicability of ACMTF-R is therefore validated by applying it to this multi-omics dataset. ACMTF-R successfully identifies novel mother-infant relationships associated with maternal pre-pregnancy BMI, underscoring its utility in multi-omics research. Our findings establish ACMTF-R as a versatile tool for multi-way data fusion, offering new insights into complex biological systems by integrating common, local, and distinct variation in the context of a dependent variable. Author SummaryIn recent years, biological research has been transformed by the rise of high-throughput technologies, allowing us to simultaneously measure multiple different data (genes, microbes, and metabolites) within the same subject. While these datasets hold great promise, analysing them in an integrated way remains challenging. Existing tools either focus on uncovering patterns in the data or on predicting outcomes, but rarely both. In this study, we present a new method called ACMTF-Regression (ACMTF-R), which combines these aspects. ACMTF-R helps researchers identify shared and distinct biological patterns across different datasets while also relating these patterns to specific outcomes. Using simulated data, we show that ACMTF-R can detect subtle signals that would otherwise go unnoticed. We also apply it to a real-world study of mothers and their infants, revealing how maternal obesity influences breast milk and gut microbes in the baby. Our approach provides a powerful new tool for studying complex biological systems and can be especially valuable in fields like microbiome research, metabolomics, and personalized medicine.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

van der Ploeg, G. R., White, F., Jakobsen, R. R., Westerhuis, J., Heintz-Buschart, A., Smilde, A.. 2025-07-31. ACMTF-R: supervised multi-omics data integration uncovering shared and distinct outcome-associated variation. https://doi.org/10.1101/2025.07.28.667162

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Concentration limits and localization of hydrogen peroxide in the extracellular space of solid tissues

H2O2 released to the extracellular space (ECS) regulates diverse physiological processes, yet its concentrations and spatial distribution in tissues remain poorly defined. This uncertainty hampers mechanistic understanding of redox signaling. Here, we used reaction-diffusion modeling to estimate extracellular H2O2 concentrations and transport ranges in various scenarios. Idealized analytical models were combined with numerical models incorporating localized NADPH oxidase (NOX) clusters, ECS microstructure, membrane permeability, and the thioredoxin- and GSH-dependent clearance systems. Using maximal neutrophil and NOX superoxide/H2O2 release rates, we obtained upper bounds for extracellular H2O2. Adjacent to isolated average-sized, fully active NOX2 clusters H2O2 peaked at ~540 nM at adhesion cell-cell separations, and decreased radially over ~50-100 nm. At the receptor cell surface, peak concentration decreased inversely with intercellular separation, to <5 nM at 1 m separation. Radial decrease here, for this wide separation, was over ~2.5 m. Even the former maximal extracellular concentrations induce just a minimal, highly localized oxidation of the intracellular Prdx, Trx and GSH pools. In turn, maximally activated neutrophils carry ~2000 such NOX2 clusters, inducing 10s of M peak ECS H2O2 concentrations. These cause extensive Prdx and Trx oxidation near the exposed membranes. However, the GSH-dependent system still sustains a strong transmembrane gradient if the permeation barrier remains intact, and ECS H2O2 concentrations decay to sub-M within a few m of the source cell. Extracellular H2O2 concentrations scaled linearly with source flux in all the examined conditions. These results establish stringent constraints on autocrine, juxtacrine and next-cell paracrine H2O2 signaling.

systems biology↗

LSD-pipeline: Causal Inference of miRNA Network Effects in Alzheimer's Disease

MicroRNAs (miRNAs) are implicated in Alzheimer's disease (AD), but research has focused on individual miRNAs and direct targets. Existing approaches to miRNA regulation in AD identify associations rather than causal effects, and few methods estimate multi-stage chains from miRNAs through target genes to target transcription factor (TF) cascades. We developed the LSD pipeline (LASSO-SEM-DoWhy), integrating LASSO feature selection, multi-stage structural equation modeling, and DoWhy causal inference to identify and validate miRNA causal pathways in AD. Applying LSD to six blood miRNA and brain mRNA datasets, we identified four LSD-validated miRNAs (miR-30d-5p, miR-92a-3p, miR-296-5p, miR-193a-5p) as AD biomarkers, achieving >86% ROC accuracy in an independent validation cohort. Several miRNAs with no significant direct association with AD showed significant effects when estimated through their target networks, while others significant in direct analysis were not supported at the network level, underscoring the value of network-level analysis. Extending to the TF layer revealed complete miRNA [->] targets [->] TF cascades [->] AD causal chains, with HMGA1, NKX2-3, and PRRX2 as key intermediaries. Confirmed classic pathways converge primarily on tau pathology and synaptic dysfunction. miRNA effects were largely age-independent, suggesting miRNAs act as early initiators of AD pathogenesis. Beyond AD, the LSD pipeline provides a generalizable framework for uncovering causal regulatory mechanisms in other diseases.

systems biology↗

PyKappa: Rule-based modeling in Python

Rule-based languages have proven effective for modeling systems of interacting structured entities as typically encountered in chemistry and molecular biology. We present PyKappa, a rule-based modeling package written in Python whose interpreted nature enables interactive simulation and analysis, including by agentic AI. The package seeks to broaden the base of developers by utilizing a widely known programming language and serves as an easy-to-deploy teaching tool. Using PyKappa, we conduct a case study of phase separation.

systems biology↗