bioRxiv Science⌕ Search

Biology subjects

Dash, H.

Publications and source records attributed to Dash, H..

5 recordsLinked to original sources

Single-Cell Study Designs Are Systematically Underpowered for Small-Effect Genes

Single cell differential expression analysis enables biologists to make statistical conclusions about which genes are up or downregulated in a particular cell type, between two conditions, such as those with or without a disease. However, due to biological and technical noise, these differences are hard to detect reliably. Determining if an experiment has sufficient power is often derived from simulations or small pilot samples, if done at all. Here, we use sex-biased differential expression on 1,494 donors in three brain cell types to derive empirically grounded power estimates for a variety of experimental setups. Our work reveals a substantial lack of power for reliably detecting the small effects in the range that many studies report, even in experiments containing 600 donors. When reducing the astrocyte data to the poor sequencing characteristics of microglia, over half the differentially expressed genes (DEGs) detected in the full set were lost, highlighting the damage caused by insufficient sequencing depth and cell counts. Despite standard thresholds of adjusted p-value with multiple testing correction, only the top quartile of significant genes were reproducible. Furthermore, from fitting a predictive model to the empirical outcomes, we find cell count and expression levels of genes to be a strong determinant of power. As such, we advocate for future studies to employ cell type enrichment and deeper sequencing, especially for rarer populations like microglia, emphasise the importance of powering experiments of this type when seeking robust findings, and suggest stricter significance thresholds for future discoveries.

genomics↗

A roadmap for enriching rare cell populations from human post-mortem brain, demonstrated by 100% microglial purity.

Understanding brain disease requires studying the specific cell types that drive each condition: microglia in Alzheimers, dopaminergic neurons in Parkinsons, motor neurons in ALS, and many others. For most of these rare populations, protocols to isolate them from human post-mortem brain at the required purity have never been developed. The shared hurdles are reliable nuclear markers, compatible fluorescent dyes, antibody host-species cross-reactivity, and achieving pure rather than merely enriched populations. Here we present a generalisable roadmap for developing fluorescence-activated nuclear sorting (FANS) protocols for rare cell populations in human brain. We demonstrate it on microglia, reaching 100% purity in cortex and 98% in cerebellum, with the cortex sort simultaneously yielding astrocytes at 93% purity. The roadmap tackles each hurdle including an on-bench antibody labelling step that expands fluorescence channels without cross-reactivity problem and an affordable single-nucleus RNA-seq validation workflow, giving labs a pathway for enriching any rare cell type reliably.

neuroscience↗

Decades of Parkinson's disease neuropathology yield a sparse and underpowered map of neuronal vulnerability: a systematic review and meta-analysis

Selective vulnerability is widely assumed in Parkinson's disease (PD), but whether dopaminergic neurons of the substantia nigra are uniquely vulnerable has not been established. Explanations of neuronal degeneration in Parkinson's disease often centre on the distinctive properties of substantia nigra dopaminergic neurons. Whether these properties are necessary for substantial neuronal loss remains unclear. We preregistered a systematic review of 166 post-mortem case-control studies published between 1963 and 2025 and synthesised neuronal counts and densities from 152 studies using a multilevel meta-analysis. After six decades, only 18% of countable brain atlas labels had been examined, and most populations were represented by a single study. Substantial loss beyond dopaminergic and classically pigmented populations shows that neither dopaminergic identity nor neuromelanin is necessary for marked degeneration, challenging key tenets of selective vulnerability in PD. Comparisons across other anatomical and physiological features remain limited by sparse sampling and uncertainty. We identify less-studied populations with substantial estimated loss and estimate the additional sampling needed to improve precision under specified assumptions. These findings direct replication and comparative counting towards uncertainties that limit explanations of neuronal loss across affected and potentially spared populations.

neuroscience↗

MotifPeeker: R package for benchmarking epigenomic profiling methods using motif enrichment as a key metric

MotifPeeker benchmarks epigenomic profiling methods targeting transcription factors (TFs) where no "gold standard" reference exists, using motif enrichment as a key metric. With minimal input, users can analyse their data in a single function and receive an intuitive HTML report. Availability and ImplementationMotifPeeker is available on Bioconductor ([≥] v3.21) at https://bioconductor.org/packages/MotifPeeker. The complete source code is available on GitHub at https://github.com/neurogenomics/MotifPeeker, with full documentation provided at https://neurogenomics.github.io/MotifPeeker. Additionally, the MotifPeeker Docker image is hosted on GitHub at https://github.com/neurogenomics/MotifPeeker/pkgs/container/motifpeeker.

bioinformatics↗

Rational design of peak calling parameters for TIP-seq based on pA-Tn5 insertion patterns improves predictive power

Epigenomic profiling provides insights into the regulatory mechanisms that govern gene expression. At a fundamental level, these mechanisms are determined by proteins that bind the DNA or modify the chromatin. Techniques such as ChIP-seq and CUT&Tag have been instrumental in mapping the binding sites of such proteins across the genome. Recent advances have led to the development of TIP-seq, a highly sensitive method devised to increase the number of unique reads per sample. Its design results in novel library features, which have not yet been explored with comparative analytics. Through the extensive assessment of bioinformatics tools and parameters we have developed an analysis pipeline that is ideally suited for TIP-seq data, including linear deduplication, read prioritisation and read shifting. Using transcription factor binding profiles (TFs), we show that our optimised pipeline greatly reduces the width of peaks to below 50% and more precisely aligns the peak summit with known motifs. A tutorial of the optimised peak calling is available on GitHub at https://github.com/neurogenomics/peak_calling_tutorial.git. Our methodological advancement substantially improves TIP-seq data quality, and the thoughtful design of analysis parameters is widely applicable to all pA-Tn5 based profiling assays.

bioinformatics↗