bioRxiv ScienceSearch

Biology subjects

Wenric, S.

Publications and source records attributed to Wenric, S..

2 recordsLinked to original sources

Using supervised learning methods for gene selection in RNA-Seq case-control studies

Whole transcriptome studies typically yield large amounts of data, with expression values for all genes or transcripts of the genome. The search for genes of interest in a particular study setting can thus be a daunting task, usually relying on automated computational methods. Moreover, most biological questions imply that such a search should be performed in a multivariate setting, to take into account the inter-genes relationships.\n\nDifferential expression analysis commonly yields large lists of genes deemed significant, even after adjustment for multiple testing, making the subsequent study possibilities extensive.\n\nHere, we explore the use of supervised learning methods to rank large ensembles of genes defined by their expression values measured with RNA-Seq in a typical 2 classes sample set. First, we use one of the variable importance measures generated by the random forests classification algorithm as a metric to rank genes. Second, we define the EPS (extreme pseudo-samples) pipeline, making use of VAEs (Variational Autoencoders) and regressors to extract a ranking of genes while leveraging the feature space of both virtual and comparable samples.\n\nWe show that, on 12 cancer RNA-Seq data sets ranging from 323 to 1210 samples, using either a random forests based gene selection method or the EPS pipeline outperforms differential expression analysis for 9 and 8 out of the 12 datasets respectively, in terms of identifying subsets of genes associated with survival.\n\nThese results demonstrate the potential of supervised learning-based gene selection methods in RNA-Seq studies and highlight the need to use such multivariate gene selection methods alongside the widely used differential expression analysis.

bioinformatics

Transcriptome wide analysis of natural antisense transcripts shows their potential role in breast cancer.

Non-coding RNAs (ncRNA) represent at least 1/5 of the mammalian transcript amount, and about 90% of the genome length is actively transcribed. Many ncRNAs have been demonstrated to play a role in cancer. Among them, natural antisense transcripts (NAT) are RNA sequences which are complementary and overlapping to those of protein-coding transcripts (PCT). NATs were punctually described as regulating gene expression, and are expected to act more frequently in cis than other ncRNAs that commonly function in trans. In this work, 22 breast cancers expressing estrogen receptors and their paired healthy tissues were analyzed by strand-specific RNA sequencing. To highlight the potential role of NATs in gene regulations occurring in breast cancer, three different gene extraction methods were used: differential expression analysis of NATs between tumor and healthy tissues, differential correlation analysis of paired NAT/PCT between tumor and healthy tissues, and NAT/PCT read count ratio variation between tumor and healthy tissues. Each of these methods yielded lists of NAT/PCT pairs that were demonstrated to be enriched in survival-associated genes on an independent cohort (TCGA). This work allows to highlight NAT lists that display a strong potential to affect the expression of genes involved in the breast cancer pathology.

genomics