bioRxiv · 10.1101/2023.08.02.551637
Compound models and Pearson residuals for normalization of single-cell RNA-seq data without UMIs
Abstract
Recent work employed Pearson residuals from Poisson or negative binomial models to normalize UMI data. To extend this approach to non-UMI data, we model the additional amplification step with a compound distribution: we assume that sequenced RNA molecules follow a negative binomial distribution, and are then replicated following an amplification distribution. We show how this model leads to compound Pearson residuals, which yield meaningful gene selection and embeddings of Smart-seq2 datasets. Further, we suggest that amplification distributions across several sequencing protocols can be described by a broken power law. The resulting compound model captures previously unexplained overdispersion and zero-inflation patterns in non-UMI data.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lause, J., Ziegenhain, C., Hartmanis, L., Berens, P., Kobak, D.. 2023-08-05. Compound models and Pearson residuals for normalization of single-cell RNA-seq data without UMIs. https://doi.org/10.1101/2023.08.02.551637
Cite the original work for its findings. Save a collection to share your selection of sources.