bioRxiv Science⌕ Search

Biology subjects

Lazecka, M.

Publications and source records attributed to Lazecka, M..

2 recordsLinked to original sources

Modelling interpretable patient-level representationsfrom structured and simple multimodal data

Patient cohort profiling increasingly includes structured views for multiple modalities, such as single-cell RNA sequencing, spatial transcriptomics or proteomics, and histology, each providing multiple subobservations per patient, including single cells, spatial spots or patches. To model such data along with simple patient-level views, current multimodal integration methods typically rely on separately precomputed summaries and fail to fully leverage information in structured views. Here we present FACTMx, a variational framework that jointly models structured and simple views to learn interpretable patient-level representations. FACTMx couples latent patient factors with subobservation clustering and per-patient component proportions, enabling direct interpretation and downstream association analyses. The framework supports different structured-view mixture assumptions, including topic- and Gaussian-structured data, while retaining modular encoder-decoder parameterisations. In simulations spanning sparse and dense dependencies and multiple noise regimes, FACTMx improved reconstruction, integration and recovery of structured components relative to previous methods. Applied to non-small cell lung cancer cohorts, FACTMx captured survival-associated latent signals linked to immune microenvironments, gene expression pathways and spatially coherent histological patterns. In a longitudinal coronary syndrome cohort, FACTMx highlighted an outcome-associated axis connected to ejection-fraction change, immune cell states, soluble mediators and cardiac injury markers. These results support joint structured-simple modelling for interpretable multimodal patient stratification.

bioinformatics↗

Cross-View Latent Integration via Nonparametric Gamma Shrinkage Factor Analysis

Factor analysis is a dominant paradigm for multi-omic heterogeneous data, but is challenged by partially redundant signals and noise across views and by an unknown true number of factors. We present CLING, an unsupervised multi-view factor model with hierarchical Bayesian sparsity priors: a product-of-Gammas prior inducing cumulative column-wise shrinkage (increasing with factor index) coupled with a Gamma-Gamma local-precision hierarchy on loadings yielding heavy-tailed marginals. This pairing enables automatic factor selection by adaptively deactivating unsupported factors while retaining active ones during inference, and induces selective sparsity that allows salient loadings to escape shrinkage while collapsing negligible ones. As a fully conjugate hierarchical model, CLING admits a scalable variational inference algorithm for multi-view data. Across synthetic benchmarks and multiomics datasets, CLING recovers more accurate factors and more informative loadings while explaining at least as much variance as competitive multi-view baselines; on glioblastoma gene expression and DNA methylation data, CLING identifies pathways linked to tumor subtype and patient age. Source code: https://github.com/szczurek-lab/CLING.

bioinformatics↗