bioRxiv Science⌕ Search

Biology subjects

Breckels, L.

Publications and source records attributed to Breckels, L..

2 recordsLinked to original sources

Robust analysis of comparative subcellular omics with complex designs

Subcellular omics technologies now allow us to obtain insights into steady-state localisation and re-localisation of biomolecules in high-throughput. However, robust analysis of these experiments can be slow and challenging. Here, we show that existing approaches to differential localisation fall into two classes, marker-dependent and marker-free, that fail for fundamentally different statistical reasons, with failure modes that are not completely overlapping. We exploit this observation in SANDLE (statistical analysis of differential localisation experiments), which combines a marker-dependent generative model of subcellular niches with a marker-free regression test for changes in fractionation profiles. This dual strategy provides better control of false positives, up to 100-fold reduction in analysis time, and accommodates more complex experimental designs than existing methods. We demonstrate SANDLEs versatility across a wide range of subcellular omics experiments, including drug-treatment responses, cross-species and life-cycle comparisons, post-translationally modified proteoforms, and RNA re-localisation. Coanalysing transcriptome and proteome dynamics during the unfolded protein response, we further reveal lncRNAs whose steady-state localisation is condition-dependent. By addressing key methodological limitations, SANDLE enables broad applications of subcellular omics from fundamental biology to clinical research.

biochemistry↗

Semi-supervised Bayesian integration of multiple spatial proteomics datasets

The subcellular localisation of proteins is a key determinant of their function. High-throughput analyses of these localisations can be performed using mass spectrometry-based spatial proteomics, which enables us to examine the localisation and relocalisation of proteins. Furthermore, complementary data sources can provide additional sources of functional or localisation information. Examples include protein annotations and other high-throughput omic assays. Integrating these modalities can provide new insights as well as additional confidence in results, but existing approaches for integrative analyses of spatial proteomics datasets are limited in the types of data they can integrate and do not quantify uncertainty. Here we propose a semi-supervised Bayesian approach to integrate spatial proteomics datasets with other data sources, to improve the inference of protein sub-cellular localisation. We demonstrate our approach outperforms other transfer-learning methods and has greater flexibility in the data it can model. To demonstrate the flexibility of our approach, we apply our method to integrate spatial proteomics data generated for the parasite Toxoplasma gondii with time-course gene expression data generated over its cell cycle. Our findings suggest that proteins linked to invasion organelles are associated with expression programs that peak at the end of the first cell-cycle. Furthermore, this integrative analysis divides the dense granule proteins into heterogeneous populations suggestive of potentially different functions. Our method is disseminated via the mdir R package available on the lead authors Github. Author summaryProteins are located in subcellular environments to ensure that they are near their interaction partners and occur in the correct biochemical environment to function. Where a protein is located can be determined from a number of data sources. To integrate diverse datasets together we develop an integrative Bayesian model to combine the information from several datasets in a principled manner. We learn how similar the dataset are as part of the modelling process and demonstrate the benefits of integrating mass-spectrometry based spatial proteomics data with timecourse gene-expression datasets.

systems biology↗