bioRxiv ScienceSearch

Biology subjects

Rachid Zaim, S.

Publications and source records attributed to Rachid Zaim, S..

2 recordsLinked to original sources

Evaluating single-subject study methods for personal transcriptomic interpretations to advance precision medicine

BackgroundGene expression profiling has benefited medicine by providing clinically relevant insights at the molecular candidate and systems levels. However, to adopt a more precision approach that integrates individual variability including omics data into risk assessments, diagnoses, and therapeutic decision making, whole transcriptome expression analysis requires methodological advancements. One need is for users to confidently be able to make individual-level inferences from whole transcriptome data. We propose that biological replicates in isogenic conditions can provide a framework for testing differentially expressed genes (DEGs) in a single subject (ss) in absence of an appropriate external reference standard or replicates.\n\nMethodsEight ss methods for identifying genes with differential expression (NOISeq, DEGseq, edgeR, mixture model, DESeq, DESeq2, iDEG, and ensemble) were compared in Yeast (parental line versus snf2 deletion mutant; n=42/condition) and MCF7 breast-cancer cell (baseline and stimulated with estradiol; n=7/condition) RNA-Seq datasets where replicate analysis was used to build reference standards from NOISeq, DEGseq, edgeR, DESeq, DESeq2. Each dataset was randomly partitioned so that approximately two-thirds of the paired samples were used to construct reference standards and the remainder were treated separately as single-subject sample pairs and DEGs were assayed using ss methods. Receiver-operator characteristic (ROC) and precision-recall plots were determined for all ss methods against each RSs in both datasets (525 combinations).\n\nResultsConsistent with prior analyses of these data, ~50% and ~15% DEGs were respectively obtained in Yeast and MCF7 reference standard datasets regardless of the analytical method. NOISeq, edgeR and DESeq were the most concordant and robust methods for creating a reference standard. Single-subject versions of NOISeq, DEGseq, and an ensemble learner achieved the best median ROC-area-under-the-curve to compare two transcriptomes without replicates regardless of the type of reference standard (>90% in Yeast, >0.75 in MCF7).\n\nConclusionBetter and more consistent accuracies are obtained by an ensemble method applied to singlesubject studies across different conditions. In addition, distinct specific sing-subject methods perform better according to different proportions of DEGs. Single-subject methods for identifying DEGs from paired samples need improvement, as no method performs with both precision>90% and recall>90%. http://www.lussiergroup.org/publications/EnsembleBiomarker

bioinformatics

iDEG: A single-subject method for assessing gene differential expression from two transcriptomes of an individual

BackgroundAccurate profiling of gene expression in a single subject has the potential to be a powerful precision medicine tool, useful for unveiling individual disease mechanisms and responses. However, expression analysis tools for RNA-sequencing (RNA-Seq) data require replicate samples to estimate gene-wise data variability and make inferences, which is costly and not easily obtainable in clinical practice. Strategies to implement DEGSeq, DESeq, and edgeR for comparing two conditions without replicates (TCWR) have been proposed without evaluation, while NOISeq-sim was validated in a restricted way using qPCR on 400 transcripts. These methods impose restrictive assumptions in TCWR limiting inferential opportunities.\n\nMethodsWe propose a new method that borrows information across different genes from the same individual using a partitioned window to strategically bypass the requirement of replicates per condition. We termed this method \"iDEG\", which identifies individualized Differentially Expressed Genes in a single subject sampled under two conditions without replicates, i.e., a baseline sample (unaffected tissue) vs. a case sample (tumor). iDEG transforms RNA-Seq data such that, under the null hypothesis, differences of transformed expression counts follow a distribution and variance calculated across a local partition of related transcripts at baseline expression. This transformation enables modeling genes with a two-group mixture model from which the probability of differential expression for each gene is then estimated by an empirical Bayes approach with a local false discovery rate control. To compare the performance of iDEG to other methods applied to TCWR, we conducted simulations assuming a Negative Binomial distribution with varying dispersion parameters and percentages of differentially expressed genes (DEGs).\n\nResultsOur extensive simulation studies demonstrate that iDEGs F1 accuracy scores better than the other methods at 5% 90% and recall>75% and low false positive rate (<1%) in most conditions.\n\nConclusionThe partitioned window strategy provides a novel and accurate way to borrow information across genes locally and would probably increase the accuracy of all relevant methods.

bioinformatics