bioRxiv Science⌕ Search

Biology subjects

Biederstedt, E.

Publications and source records attributed to Biederstedt, E..

4 recordsLinked to original sources

Scaling cross-tissue single-cell annotation models

Identifying cellular identities (both novel and well-studied) is one of the key use cases in single-cell transcriptomics. While supervised machine learning has been leveraged to automate cell annotation predictions for some time, there has been relatively little progress both in scaling neural networks to large data sets and in constructing models that generalize well across diverse tissues and biological contexts up to whole organisms. Here, we propose scTab, an automated, feature-attention-based cell type prediction model specific to tabular data, and train it using a novel data augmentation scheme across a large corpus of single-cell RNA-seq observations (22.2 million human cells in total). In addition, scTab leverages deep ensembles for uncertainty quantification. Moreover, we account for ontological relationships between labels in the model evaluation to accommodate for differences in annotation granularity across datasets. On this large-scale corpus, we show that cross-tissue annotation requires nonlinear models and that the performance of scTab scales in terms of training dataset size as well as model size - demonstrating the advantage of scTab over current state-of-the-art linear models in this context. Additionally, we show that the proposed data augmentation schema improves model generalization. In summary, we introduce a de novo cell type prediction model for single-cell RNA-seq data that can be trained across a large-scale collection of curated datasets from a diverse selection of human tissues and demonstrate the benefits of using deep learning methods in this paradigm. Our codebase, training data, and model checkpoints are publicly available at https://github.com/theislab/scTab to further enable rigorous benchmarks of foundation models for single-cell RNA-seq data.

bioinformatics↗

Gene panel selection for targeted spatial transcriptomics

Targeted spatial transcriptomics hold particular promise in analysis of complex tissues. Most such methods, however, measure only a limited panel of transcripts, which need to be selected in advance to inform on the cell types or processes being studied. A limitation of existing gene selection methods is that they rely on scRNA-seq data, ignoring platform effects between technologies. Here we describe gpsFISH, a computational method to perform gene selection through optimizing detection of known cell types. By modeling and adjusting for platform effects, gpsFISH outperforms other methods. Furthermore, gpsFISH can incorporate cell type hierarchies and custom gene preferences to accommodate diverse design requirements.

bioinformatics↗

Tensor decomposition reveals coordinated multicellular patterns of transcriptional variation that distinguish and stratify disease individuals

Tissue- and organism-level biological processes often involve coordinated action of multiple distinct cell types. Current computational methods for the analysis of single-cell RNA-sequencing (scRNA-seq) data, however, are not designed to capture co-variation of cell states across samples, in part due to the low number of biological samples in most scRNA-seq datasets. Recent advances in sample multiplexing have enabled population-scale scRNA-seq measurements of tens to hundreds of samples. To take advantage of such datasets, here we introduce a computational approach called single-cell Interpretable Tensor Decomposition (scITD). This method extracts "multicellular gene expression patterns" that capture how sample-specific expression states of a cell type are correlated with the expression states of other cell types. Such multicellular patterns can reveal molecular mechanisms underlying coordinated changes of different cell types within the tissue, and can be used to stratify individuals in a clinically-relevant and reproducible manner. We first validated the performance of scITD using in vitro experimental data and simulations. We then applied scITD to scRNA-seq data on peripheral blood mononuclear cells (PBMCs) from 115 patients with systemic lupus erythematosus and 56 healthy controls. We recapitulated a well-established pan-cell-type signature of interferon-signaling that was associated with the presence of anti-dsDNA autoantibodies and a disease activity index. We further identified a novel multicellular pattern linked to nephritis, which was characterized by an expansion of activated memory B cells along with helper T cell activation. Our approach also sheds light on ligand-receptor interactions potentially mediating these multicellular patterns. As validation, we demonstrated that these expression patterns also stratified donors from a pediatric SLE dataset by the same phenotypic attributes. Lastly, we found the interferon multicellular pattern and others to be conserved in a COVID-19 dataset, pointing to the presence of both general and disease-specific patterns of inter-individual immune variation. Overall, scITD is a flexible method for exploring co-variation of cell states in multi-sample single-cell datasets, which can yield new insights into complex non-cell-autonomous dependencies that define and stratify disease.

bioinformatics↗

Haplotype-enhanced inference of somatic copy number profiles from single-cell transcriptomes

Genome instability and aberrant alterations of transcriptional programs both play important roles in cancer. However, their relationship and relative contribution to tumor evolution and therapy resistance are not well-understood. Single-cell RNA sequencing (scRNA-seq) has the potential to investigate both genetic and non-genetic sources of tumor heterogeneity in a single assay. Here we present a computational method, Numbat, that integrates haplotype information obtained from population-based phasing with allele and expression signals to enhance detection of CNVs from scRNA-seq data. To resolve tumor clonal architecture, Numbat exploits the evolutionary relationships between subclones to iteratively infer the single-cell copy number profiles and tumor clonal phylogeny. Analyzing 21 tumor samples composed of multiple myeloma, breast, and thyroid cancers, we show that Numbat can accurately reconstruct the tumor copy number profile and precisely identify malignant cells in the tumor microenvironment. We uncover additional subclonal complexity contributed by allele-specific alterations, and identify genetic subpopulations with transcriptional signatures relevant to tumor progression and therapy resistance. We hope that the increased power to characterize genomic aberrations and tumor subclonal phylogenies provided by Numbat will help delineate contributions of genetic and non-genetic mechanisms in cancer.

bioinformatics↗