bioRxiv Science⌕ Search

Biology subjects

Gaetano, L.

Publications and source records attributed to Gaetano, L..

2 recordsLinked to original sources

MIDFA: Scalable Bayesian Factor Analysis for Mixed and Incomplete Data

Probabilistic latent variable models are a powerful tool for uncovering structure in high-dimensional datasets, particularly in biomedical applications. The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high dimensionality, and structured missingness. Existing approaches address some of these issues, but few provide a unified and scalable framework for handling them simultaneously. Here, we propose a scalable Bayesian factor analysis framework designed to address these challenges. Our method combines a semi-parametric Gaussian copula model with a continuous spike-and-slab prior to induce sparse and interpretable factor loadings. The number of latent dimensions is learned nonparametrically from the data using an Indian buffet process prior. For model fitting, we develop an expectation-maximisation algorithm that naturally accommodates missing data. We validate the proposed method through comprehensive simulation studies. In addition, we showcase the proposed model using the Novartis-Oxford Multiple Sclerosis dataset in two ways. First, we identify latent dimensions shared across MS clinical and neuroimaging variables, characterising disease structure while demonstrating the models ability to handle multiple data types and structured missingness. Second, we use the model for dimensionality reduction of structural MRI data, extracting features for downstream analysis that go beyond traditional whole-brain summary statistics. In those applications, our method identifies sparse latent structures and provides insights beyond those obtained from traditional approaches.

neuroscience↗

BARTharm: MRI Harmonization Using Image Quality Metrics and Bayesian Non-parametric

Image derived phenotypes (IDPs) harmonization from Magnetic Resonance Imaging (MRI) data is essential for reducing scanner-induced, non-biological variability and enabling accurate multi-site analysis. Existing methods like ComBat, while widely used, rely on linear assumptions and explicit scanner IDs - limitations that reduce their effectiveness in real-world scenarios involving complex scanner effects, non-linear biological variation, or anonymized data. We introduce BARTharm, a novel harmonization framework that uses Image Quality Metrics (IQMs) instead of Scanner IDs and models scanner and biological effects separately using Bayesian Additive Regression Trees (BART), allowing for flexible, data-driven adjustment of IDPs. Through extensive simulation studies, we demonstrate that IQMs provide a more informative and flexible representation of scanner-related variation than categorical Scanner IDs, enabling more accurate removal of non-biological effects. Leveraging this and its ability to model complex relationships, BARTharm, consistently outperforms ComBat across a range of challenging scenarios, including model misspecification and confounded scanner-biological relationships. Applied to real-world datasets, BARTharm successfully removes scanner-induced bias while preserving meaningful biological signals, resulting in stronger, more reliable associations with clinical outcomes. Overall, we find that BARTharm is a robust, data-driven improvement over traditional harmonization approaches, particularly suited for modern, large-scale neuroimaging studies.

neuroscience↗