bioRxiv Science⌕ Search

Biology subjects

Hoerbst, F.

Publications and source records attributed to Hoerbst, F..

3 recordsLinked to original sources

What is a differentially expressed gene?

The concept of Differentially Expressed Genes (DEGs) is central to RNA-Seq studies, yet their identification suffers from reproducibility issues. This is largely a consequence of the inherent biological and technical variation that cannot be captured with small numbers of replicates. When thresholds for p-values and log2 fold changes are introduced, this variability can propagate an incomplete description of the data, leading to differing interpretations. Here, we compare traditional binary DEG classification with a rank-based method, grounded in Bayesian statistics, using a published yeast dataset comprising over 40 replicates. This analysis reveals how the choice of thresholds and number of replicates results in discrepancies between studies and potentially interesting genes being overlooked. Furthermore, by comparing wild-type with wild-type samples, we show how variability in gene expression can be mistaken for differential expression. Evaluating current practices for navigating the accuracy-error trade-off in the search for differentially expressed genes leads us to advocate rank-based methods and Bayesian statistics to mitigate the limitations of binary classifications and communicate uncertainty.

bioinformatics↗

A Bayesian framework for ranking genes based on their statistical evidence for differential expression

Advances in sequencing technologies have revolutionised our ability to capture the complete RNA profile in tissue samples, allowing for comparative analyses of RNA levels between developmental processes, environmental responses, or treatments. However, quantifying changes in gene expression remains challenging, given inherent biological variability and technological limitations. To address this, we introduce a Bayesian framework for differential gene expression (DGE) analysis. Our framework unifies and streamlines a complex analysis, typically involving parameter estimations and multiple statistical tests, into a concise mathematical equation. This allows statistical evidence for differential expression to be computed rapidly and transparently. We show how this approach can be used to evaluate variabilty of individual genes between replicates. A comparison of our framework with existing tools revealed differences that can be explained by commonly employed thresh-olds in other packages. This motivated us to explore ranking genes based on their statistical evidence as opposed to a binary classification as DEGs. Our analysis leads us to advocate the use of Bayes factors within a rank-based approach. This framework offers enhanced computational efficiency and delivers a transparent way to analyse, interpret and communicate DGE results.

bioinformatics↗

Re-analysis of mobile mRNA datasets highlights challenges in the detection of mobile transcripts from short-read RNA-Seq data

Short-read RNA-Seq analyses of grafted plants have led to the proposal that large numbers of mRNAs move over long distances between plant tissues, acting as potential signals. The detection of transported transcripts by RNA-Seq is both experimentally and computationally challenging, requiring successful grafting, delicate harvesting, rigorous contamination controls and data processing approaches that can identify rare events in inherently noisy data. Here, we perform a meta-analysis of existing datasets and examine the associated bioinformatic pipelines. Our analysis reveals that technological noise, biological variation and incomplete genome assemblies give rise to features in the data that can distort the interpretation. Taking these considerations into account, we find that a substantial number of transcripts that are currently annotated as mobile are left without support from the available RNA-Seq data. Whilst several annotated mobile mRNAs have been validated, we cannot exclude that others may be false positives. The identified issues may also impact other RNA-Seq studies, in particular those using single nucleotide polymorphisms (SNPs) to detect variants.

plant biology↗