bioRxiv Science⌕ Search

Biology subjects

Sarwar, A.

Publications and source records attributed to Sarwar, A..

2 recordsLinked to original sources

Detecting cell segmentation errors using doublet methods

In spatial transcriptomics, cell segmentation is used to draw boundaries around cells. Molecules located within a cell's boundary are assigned to it, making its gene expression profile dependent on segmentation accuracy. To identify potentially problematic cells, studies increasingly use doublet detection methods, a class of algorithms developed to recognize molecular admixture from two cells in scRNA-seq. To evaluate their suitability in spatial data, we model cell segmentation errors as a continuum of partial molecular admixture between neighboring cells, generated by varying the loss of a cell's own transcripts and the gain of transcripts from its neighbor. Evaluating 8 doublet methods across 16 spatial datasets, we characterize the conditions under which they perform well and identify their failure modes. Detection improves with increasing molecular admixture and transcriptional dissimilarity between neighboring cells, with cxds2, scDblFinder.score, and a simple baseline (unique_genes) performing best. These patterns are consistent across datasets spanning different gene panels, technological platforms, staining techniques, segmentation algorithms, and error frequencies. In unperturbed spatial data, elevated doublet scores localize to regions consistent with segmentation problems, suggesting that they can help prioritize cells or regions for inspection, segmentation refinement, or transcript reassignment. Our results establish when doublet methods can provide useful quality control signals for cell segmentation, supporting their use in spatial transcriptomics.

bioinformatics↗

Cross-expression analysis reveals patterns of coordinated gene expression in spatial transcriptomics

Spatial transcriptomics promises to transform our understanding of tissue biology by molecularly profiling individual cells in situ. A fundamental question they allow us to ask is how nearby cells orchestrate their gene expression. Rather than focus on how these cells (samples) communicate with each other, we reframe the problem to investigate how genes (features) coordinate their expression between neighboring cells. To study these phenomena - called cross-expression - we compare all genes to find pairs that coordinate their expression between adjacent cells, thereby avoiding curating gene lists or annotating cell types. Our end-to-end method recovers ligand-receptor pairs as cross-expressing genes and finds gene combinations that mark anatomical regions, complementing marker gene-based region annotation. Leveraging the overlapping genes across different panels, we use multiple atlas-scale adult mouse brain datasets (~25 million cells, 695 samples, 8 technologies) to create an integrated, meta-analytic cross-expression network, whose communities are enriched in spatial processes such as synaptic signaling and G protein coupled receptor activity. Highlighting cross-expressions biological utility, our network shows that genes Drd1 and Gpr6, which are individually implicated in Parkinsons disease (PD) and are being pursued as therapeutic targets, are cross-expressed within the striatum, hinting at their joint role in PD pathophysiology. We provide an efficient R package (https://github.com/gillislab/CrossExpression/) to computationally analyze and visually explore cross-expression patterns, which allow us to better understand how genes coordinate their expression in space to perform tissue-level functions.

bioinformatics↗