bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.01.09.698732

SCALPEL: A pipeline for processing large-scale spatial transcriptomics data

Abstract

Spatial transcriptomics enables the precise mapping of gene expression patterns within tissue architecture, offering unprecedented insights into cellular interactions, tissue heterogeneity, and disease pathology that are unattainable with traditional transcriptomic approaches. We present a tool for processing spatial transcriptomics data, SCALPEL (Spatial Cell Analysis, Labeling, Processing, and Expression Linking). SCALPEL is specifically designed to support the analysis of large, atlas-level datasets. Our new workflow features advanced 3D segmentation optimized for dense and heterogeneous tissues, refined filtering criteria, and transcriptome-based doublet detection to remove low-quality or artifactual cells. Cell type label transfer from existing taxonomies is further improved through updated filtering thresholds. Spatial domain detection is incorporated to capture local transcriptomic organization, and tissue sections are registered to the Allen Mouse Brain Common Coordinate Framework version 3 (CCFv3) for precise anatomical alignment. Genome-wide expression imputation from single-cell RNA-sequencing (scRNAseq) further enriches the dataset. Crucially, we benchmark the performance of this updated pipeline against a previously published version of our whole-mouse-brain (WMB) dataset (Yao et al., 2023b), demonstrating substantial improvements in cell number, expression profile clarity, and spatial registration. These advances provide a robust foundation for downstream spatial analyses and set a new standard for large-scale spatial transcriptomics studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kunst, M., Ching, L., Quon, J., Mathieu, R., Hewitt, M., Seeman, S., Ayala, A., Gelfand, E., Long, B., Martin, N., Nagra, J., Olsen, P., Oyama, A., Valera, N., Pagen, C., Sunkin, S., Ariza, J., Smith, K., McMillen, D., Zeng, H., Waters, J.. 2026-01-12. SCALPEL: A pipeline for processing large-scale spatial transcriptomics data. https://doi.org/10.64898/2026.01.09.698732

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Modular core network constructed from Escherichia coli transcriptome datasets using a hypergraph-based pan-network approach

We have developed a method to integrate transcriptomic coexpression networks across diverse experimental conditions within a single species. Our framework expands previous pannetwork approaches into a hypergraph based pannetwork. It first identifies coexpressed gene clusters within each individual dataset and extracts frequently co-expressed gene sets across multiple datasets using a frequent itemset mining algorithm. This process yields a hypergraph where each hyperedge is assigned a frequency (universality, U). We then extracted a subnetwork comprising high U hyperedges as the core network and formed modules within it. We applied our method to 106 Escherichia coli transcriptome datasets from the GEO database. The modularity of the core network peaked at a universality cutoff of 15, which was subsequently used to define it. Approximately 70% of the resulting core modules correspond to operons, and conversely, approximately 70% of all operons are covered by these core modules. We visualized the core network via an inter-modular network and analyzed core module dataset relationships using a modularity profile matrix. Based on these analyses, we successfully visualized the dynamic reorganization of the coexpression network in response to environmental changes in bacteria.

bioinformatics↗

Phase Separation Potential of Marsupial RSX RNA Reveals Convergent Evolution of X-Chromosome Inactivation Mechanisms

Background: X-chromosome inactivation (XCI) evolved independently in eutherian and marsupial mammals, where it is orchestrated by the unrelated long non-coding RNAs Xist and RSX, respectively. Xist organizes a repressive nuclear compartment through multivalent RNA-protein interactions, but whether RSX exploits similar biophysical principles remains unknown. A recent paper has identified bona fide RSX interacting proteins. Results: We integrated proteome-scale RNA-protein interaction prediction, experimental validation, phase-separation propensity analysis, functional annotation and comparative RNA-structure modelling to characterize the RSX interaction landscape. Using catRAPID, we ranked 1,168 RNA-binding proteins from the native Monodelphis domestica proteome. Predictions were significantly enriched for experimentally identified RSX interactors, with 4.85-fold enrichment among the top 50 candidates (P approximately 1.6 x 10-5), increasing to approximately eightfold for proteins shared by the experimental RSX and Xist interactomes (P approximately 2 x 10-6). Among 30 high-confidence RSX interactors, 13 were experimentally supported, 17 were previously unrecognized candidates and 17 exhibited high phase-separation propensity. The network was enriched in ribonucleoprotein granules and nuclear bodies and converged on m6A regulators and SR-family splicing factors. Comparative modelling detected no conserved secondary or tertiary architecture between RSX and Xist. Conclusions: RSX and Xist appear to have converged not through RNA sequence or global structure, but through recruitment of related, condensation-prone protein networks. These findings identify interaction-network and biophysical convergence as a potential principle of lncRNA-mediated chromosome regulation and provide testable candidates for determining whether RSX establishes a condensate-like compartment on the marsupial inactive X.

bioinformatics↗

Systematic discovery of protein kinase-like domains reveals diverse evolutionary strategies in the human oral microbiome

The protein kinase-like (PKL) superfamily regulates metabolism, biofilm formation, host interactions, and antimicrobial resistance in bacteria, although much of its diversity remains unknown. We used a structure-based pipeline that combined protein structure prediction with structural similarity searches to search 5.1 million protein sequences in the Human Oral Microbiome Database (HOMD). Together with 30 families that had been previously characterized, we identified 20 PKL families new to this genome collection: 15 entirely novel and 5 previously reported as preliminary findings. Our taxonomic analysis revealed that the families were distributed either across multiple bacterial phyla or restricted to a single species; hits spanning domains suggest either ancient origins or horizontal transfer. Three families were examined in detail and show different evolutionary pathways: SEAE1 from Segetibacter aerophilus, which is a putative lipid kinase; POGI1 from Porphyromonas gingivalis, a putative ethanolamine kinase with the PKL fold that is limited to pathogens; and SrfA-N from Haemophilus parainfluenzae, a non-catalytic scaffold that has been convergently co-opted as a tripartite toxin platform. Taken together, these examples demonstrate enzymatic specialization, adaptation specific to pathogens and structural co-option. The families identified are candidates for further mechanistic study.

bioinformatics↗