bioRxiv Science⌕ Search

Biology subjects

Vandenbulcke, S.

Publications and source records attributed to Vandenbulcke, S..

2 recordsLinked to original sources

gamdid: generalized additive models for differential distributions in single cell experiments

Single-cell proteomics (SCP) generates protein abundance measurements across hundreds to thousands of individual cells, offering unprecedented resolution to study cellular heterogeneity. However, existing differential abundance (DA) methods are limited to detecting shifts in mean expression, leaving biologically relevant differences in shape undetected. Indeed, the specific power of SCP is to identify differences between individual cells in a population, which are typically only found as shape differences rather than in mean expression. We here therefore present gamdid (generalized additive models for differential distributions), a novel statistical framework and R package for differential distribution (DD) analysis in SCP data. gamdid is based on generalized additive models (GAMs) to flexibly model heterogeneous distributions, perform inference and provide interpretable visualizations. Through semi-synthetic benchmarking on two SCP datasets, gamdid demonstrates conservative false discovery rate control and substantially outperforms competing methods for differences in shape, while achieving comparable performance for mean shifts. A spike-in case study further demonstrates the utility of gamdid and its interpretable visualization. Uniquely among DD methods, gamdid supports omnibus testing across more than two groups, with post-hoc pairwise comparisons via stagewise testing, and is specifically tailored for proteomics abundance data.

genomics↗

msqrob2TMT: robust linear mixed models for inferring differential abundant proteins in labelled experiments with arbitrarily complex design

Labelling strategies in mass spectrometry (MS)-based proteomics enable increased sample throughput by acquiring multiplexed samples in a single run. However, contemporary designs often require the acquisition of multiple runs, leading to a complex correlation structure. Addressing this correlation is key for correct statistical inference and reliable biomarker discovery. Therefore, we present msqrob2TMT, a set of mixed model-based workflows tailored toward differential abundance analysis for labelled MS-based proteomics data. Thanks to its increased flexibility, msqrob2TMT can model both sample-specific and feature-specific (e.g. peptide or protein) covariates, which unlocks the inference to experiments with arbitrarily complex designs as well as to correct explicitly for feature-specific properties. We benchmark our novel workflows against the state-of-the-art tools MSstatsTMT and DeqMS in a spike-in study. We show that our workflows are modular, more flexible and have improved performance by adopting robust ridge regression. We also found that reference channel normalization and imputation can have a deleterious impact on the statistical outcome. Finally, we demonstrate the significance of msqrob2TMT on a real-life mice study, showcasing the importance of effectively accounting for the hierarchical correlation structure in the data.

bioinformatics↗