bioRxiv Science⌕ Search

Biology subjects

Peyre, G.

Publications and source records attributed to Peyre, G..

3 recordsLinked to original sources

Paired single-cell multi-omics data integration with Mowgli

The profiling of multiple molecular layers from the same set of cells has recently become possible. There is thus a growing need for multi-view learning methods able to jointly analyze these data. We here present Multi-Omics Wasserstein inteGrative anaLysIs (Mowgli), a novel method for the integration of paired multi-omics data with any type and number of omics. Of note, Mowgli combines integrative Nonnegative Matrix Factorization (NMF) and Optimal Transport (OT), enhancing at the same time the clustering performance and interpretability of integrative NMF. We apply Mowgli to multiple paired single-cell multi-omics data profiled with 10X Multiome, CITE-seq and TEA-seq. Our in depth benchmark demonstrates that Mowglis performance is competitive with the state-of-the-art in cell clustering and superior to the state-of-the-art once considering biological interpretability. Mowgli is implemented as a Python package seamlessly integrated within the scverse ecosystem and it is available at http://github.com/cantinilab/mowgli.

bioinformatics↗

Climatic refugia in the coldest neotropical hotspot, the Andean paramo

AimThe Andean paramo is the most biodiverse high-mountain region on Earth and past glaciation dynamics during the Quaternary are greatly responsible for its plant diversification. Here, we aim at identifying potential climatic refugia since the Last Glacial Maximum (LGM) in the paramo, according to plant family, biogeographic origin, and life-form. LocationThe paramo region in the Northern Andes MethodsWe built species distribution models for 664 plant species to generate range maps under current and LGM conditions, using five General Circulation Models (GCMs). For each species and GCM, we identified potential (suitable) and potential active (likely still occupied) refugia where both current and LGM range maps overlap. We stacked and averaged the resulting refugia maps across species and GCMs to generate consensus maps for all species, plant families, biogeographic origins and life-forms. All maps were corrected for potential confounding effect due to species richness. ResultsWe found refugia to be chiefly located in the southern and central paramos of Ecuador and Peru, especially towards the paramo ecotone with lower-elevation forests. However, we found additional specific patterns according to plant family, biogeographic origin and life-form. For instance, endemics showed refugia concentrated in the northern paramos. Main conclusionsOur findings suggest that large and connected paramo areas, but also the transitional Amotape-Huancabamba zone with the Central Andes, are primordial areas for plant species refugia since the LGM. This study therefore enriches our understanding on paramo evolution and calls for future research on plant responses to future climate change.

ecology↗

Optimal Transport improves cell-cell similarity inference in single-cell omics data

The recent advent of high-throughput single-cell molecular profiling is revolutionizing biology and medicine by unveiling the diversity of cell types and states contributing to development and disease. The identification and characterization of cellular heterogeneity is typically achieved through unsupervised clustering, which crucially relies on a similarity metric. We here propose the use of Optimal Transport (OT) as a cell-cell similarity metric for single-cell omics data. OT defines distances to compare, in a geometrically faithful way, high-dimensional data represented as probability distributions. It is thus expected to better capture complex relationships between features and produce a performance improvement over state-of-the-art metrics. To speed up computations and cope with the high-dimensionality of single-cell data, we consider the entropic regularization of the classical OT distance. We then extensively benchmark OT against state-of-the-art metrics over thirteen independent datasets, including simulated, scRNA-seq, scATAC-seq and single-cell DNA methylation data. First, we test the ability of the metrics to detect the similarity between cells belonging to the same groups (e.g. cell types, cell lines of origin). Then, we apply unsupervised clustering and test the quality of the resulting clusters. In our in-depth evaluation, OT is found to improve cell-cell similarity inference and cell clustering in all simulated and real scRNA-seq data, while its performances are comparable with Pearson correlation in scATAC-seq and single-cell DNA methylation data. All our analyses are reproducible through the OT-scOmics Jupyter notebook available at https://github.com/ComputationalSystemsBiology/OT-scOmics.

systems biology↗