bioRxiv Science⌕ Search

Biology subjects

Bakulin, A.

Publications and source records attributed to Bakulin, A..

4 recordsLinked to original sources

CytoVI: Deep generative modeling of antibody-based single cell technologies

Due to their robustness, dynamic range and scalability, antibody-based single cell technologies, such as flow cytometry, mass cytometry and CITE-seq, have become an irreplaceable part of routine clinics and a powerful tool for basic research. However, their analysis is complicated by measurement noise and bias, differences between batches, technology platforms, and restricted antibody panels. This results in a limited capacity to accumulate knowledge across technologies, studies, experimental batches, or across different antibody panels. Here, we present CytoVI - a probabilistic generative model designed to address these challenges and enable statistically rigorous and integrative analysis for antibody-based single cell technologies. We show that CytoVI outperforms existing computational methods and effectively handles a variety of integration scenarios. CytoVI enables key functionalities such as generating informative cell embeddings, imputing missing measurements, differential protein expression testing, and automated annotation of cells. We applied CytoVI to generate an integrated B cell maturation atlas across 350 proteins from a set of smaller antibody panels measured by conventional mass cytometry, and identified proteins associated with immunoglobulin class-switching in healthy humans. Using a cohort of B cell non-Hodgkin lymphoma patients measured by flow cytometry, CytoVI uncovered T cell states that are associated with disease. Finally, we show that CytoVI is a robust probabilistic framework for the analysis of standard diagnostic flow cytometry antibody panels, enabling the automated detection of tumor populations and diagnoses of incoming patient samples. CytoVI facilitates accurate and automated analysis in both preclinical and clinical settings and is available as open-source software at scvi-tools.org.

bioinformatics↗

scVIVA: a probabilistic framework for representation of cells and their environments in spatial transcriptomics

Spatial transcriptomics provides a significant advance over studies of dissociated cells in that it reveals the environment in which cells reside, thus opening the way for a more complete description of their state and function. However, most current methods for embedding and discovery of cell states rely only on the cells own gene expression profile, thus raising the need for ways to account for the neighboring cells as well. Here, we introduce scVIVA, a deep generative model that leverages both cell-intrinsic and neighboring gene expression profiles to output stochastic embeddings of cell states as well as normalized gene expression profiles. We demonstrate that scVIVA produces informative fine-grained partitions of cells that reflect both their internal state and the surrounding tissue and that its generative model facilitates the testing of hypotheses of differential expression between tissue niches. We leverage these properties of scVIVA to uncover a spatially-restricted tumor-promoting endothelial population in breast cancer and niche-associated T cell states that are shared across multiple cancers. scVIVA is available as open source software within scvi-tools.org.

bioinformatics↗

Ribonanza: deep learning of RNA structure through dual crowdsourcing

Prediction of RNA structure from sequence remains an unsolved problem, and progress has been slowed by a paucity of experimental data. Here, we present Ribonanza, a dataset of chemical mapping measurements on two million diverse RNA sequences collected through Eterna and other crowdsourced initiatives. Ribonanza measurements enabled solicitation, training, and prospective evaluation of diverse deep neural networks through a Kaggle challenge, followed by distillation into a single, self-contained model called RibonanzaNet. When fine tuned on auxiliary datasets, RibonanzaNet achieves state-of-the-art performance in modeling experimental sequence dropout, RNA hydrolytic degradation, and RNA secondary structure, with implications for modeling RNA tertiary structure.

biophysics↗

Addressing biases in gene-set enrichment analysis: a case study of Alzheimer's Disease

Inferring the driving regulatory programs from comparative analysis of gene expression data is a cornerstone of systems biology. Many computational frameworks were developed to address this problem, including our iPAGE (information-theoretic Pathway Analysis of Gene Expression) toolset that uses information theory to detect non-random patterns of expression associated with given pathways or regulons1. Our recent observations, however, indicate that existing approaches are susceptible to the biases and artifacts that are inherent to most real world annotations. To address this, we have extended our information-theoretic framework to account for specific biases in biological networks using the concept of conditional information. This novel implementation, called pyPAGE, provides an unbiased way for the estimation of the activity of transcriptional and post-transcriptional regulons. To showcase pyPAGE, we performed a comprehensive analysis of regulatory perturbations that underlie the molecular etiology of Alzheimers disease (AD). pyPAGE successfully recapitulated several known AD-associated gene expression programs. We also discovered several additional regulons whose differential activity is significantly associated with AD. We further explored how these regulators relate to pathological processes in AD through cell-type specific analysis of single cell gene expression datasets.

bioinformatics↗