bioRxiv Science⌕ Search

Biology subjects

Kalson, L.

Publications and source records attributed to Kalson, L..

2 recordsLinked to original sources

PACMOS: an R package for Projection And Classification of Multi-Omic Samples

MotivationIntegrated multi-omic analyses have transformed our understanding of cancer biology, giving rise to data-driven molecular classifications that capture disease heterogeneity beyond conventional histopathology. Among these approaches, multi-omic factor analysis (MOFA), a multimodal extension of principal component analysis, has been widely used to identify sources of molecular variation across omic layers and classify samples into molecular groups. However, classifying query samples according to an existing MOFA-based classification remains challenging, as there is no validated computational method for projecting samples into pretrained MOFA latent factor spaces. ResultsWe present PACMOS, an R package that provides a generalizable approach to project query samples into pretrained MOFA latent factor spaces. We validate PACMOS using two cancer datasets with published MOFA-based classifications--lung neuroendocrine neoplasms and pleural mesothelioma--showing that PACMOS preserves the existing MOFA latent factor space while allowing query samples to be classified. Availability and implementationPACMOS is an open-source R package available at https://github.com/IARCbioinfo/PACMOS and archived on Zenodo at https://doi.org/10.5281/zenodo.20933824, along with installation instructions and a vignette. Supplementary informationSupplementary data are available in separate files. Key messagesO_LIPACMOS enables the projection of query cancer samples into pretrained multi-omic latent factor spaces. C_LIO_LIThe package supports both continuous and discrete classifications. C_LIO_LIPACMOS provides a reproducible, per-sample workflow implemented in an R package. C_LIO_LIPACMOS demonstrates robust performance on pleural mesothelioma and lung neuroendocrine tumor datasets. C_LI

bioinformatics↗

A multi-omic, spatial, and whole-slide image dataset of lung neuroendocrine tumours from the lungNENomics cohort

Lung neuroendocrine tumours (lung NETs) are rare neoplasms comprising approximately 2% of lung cancers. Recent studies have identified distinct molecular groups based on transcriptome and methylome data, but genomic and morphological features remain underexplored due to limited whole-genome and imaging data. We have generated the largest multi-omic dataset of lung NETs to date (201 participants, for a total of n = 294 tumours), including RNA sequencing, EPIC 850K methylation arrays, and whole-genome sequencing. This multiomic dataset also include multi-regional whole-genome sequencing for 41 participants, allowing for the quantification of intra-tumoural heterogeneity. We additionally generated spatial proteomics (64 participants), spatial transcriptomics (4 participants) and whole-slide histopathology images for 212 cases. This dataset enables a comprehensive characterization of lung NET molecular groups and the identification of group-specific morphological features using deep learning algorithms. All quality control analyses, processed data, and scripts are provided to ensure reproducibility. This dataset is available as a basis for further molecular and morphological analysis of lung NETs, and for future research on multi-scale integration.

bioinformatics↗