bioRxiv ScienceSearch

Biology subjects

Thas, O.

Publications and source records attributed to Thas, O..

3 recordsLinked to original sources

A unified framework for unconstrained and constrained ordination of microbiome read count data

Explorative visualization techniques provide a first summary of microbiome read count datasets through dimension reduction. A plethora of dimension reduction methods exists, but many of them focus primarily on sample ordination, failing to elucidate the role of the bacterial species. Moreover, implicit but often unrealistic assumptions underlying these methods fail to account for overdispersion and differences in sequencing depth, which are two typical characteristics of sequencing data. We combine log-linear models with a dispersion estimation algorithm and flexible response function modelling into a framework for unconstrained and constrained ordination. The method allows easy filtering of technical confounders. As opposed to most existing ordination methods, the assumptions underlying the method are stated explicitly and can be verified using simple diagnostics. The combination of unconstrained and constrained ordination in the same framework is unique in the field and greatly facilitates microbiome data exploration. We illustrate the advantages of our method on simulated and real datasets, while pointing out flaws in existing methods. The algorithms for fitting and plotting are available in the R-package RCM.

microbiology

A Comprehensive Overview Of Genomic Imprinting In Breast & Its Deregulation In Cancer

Genomic imprinting, the parent-of-origin specific monoallelic expression of genes, plays an important role in growth and development. Loss of imprinting of individual genes has been found in varying cancers, yet data-analytical challenges have impeded systematic studies so far. We developed a mixture distribution model to detect monoallelically expressed loci in a genome-wide manner without the need for genotyping data, and applied the methodology on TCGA breast tissue RNA-seq data. We identified 35 putatively imprinted genes in healthy breast. In breast cancer however, HM13 was featured by significant loss of imprinting and expression upregulation, which could be linked to DNA demethylation. Other imprinted genes (25 out of 35) demonstrated consistent expression downregulation in breast cancer, which often correlated with loss of imprinting. A breast imprinted gene network, deregulated in cancer, might hence be present. In summary, our novel methodology highlights the massive deregulation of imprinting in breast cancer.

cancer biology

Differential gene expression analysis tools exhibit substandard performance for long non-coding RNA-sequencing data

BackgroundProtein-coding RNAs (mRNA) have been the primary target of most transcriptome studies in the past, but in recent years, attention has expanded to include long non-coding RNAs (lncRNA). lncRNAs are typically expressed at low levels, and are inherently highly variable. This is a fundamental challenge for differential expression (DE) analysis. In this study, the performance of 14 popular tools for testing DE in RNA-seq data along with their normalization methods is comprehensively evaluated, with a particular focus on lncRNAs and low abundant mRNAs.\n\nResultsThirteen performance metrics were used to evaluate DE tools and normalization methods using simulations and analyses of six diverse RNA-seq datasets. Non-parametric procedures are used to simulate gene expression data in such a way that realistic levels of expression and variability are preserved in the simulated data. Throughout the assessment, we kept track of the results for mRNA and lncRNA separately. All statistical models exhibited inferior performance for lncRNAs compared to mRNAs across all simulated scenarios and analysis of benchmark RNA-seq datasets. No single tool uniformly outperformed the others.\n\nConclusionOverall, the linear modeling with empirical Bayes moderation (limma) and the nonparametric approach (SAMSeq) showed best performance: good control of the false discovery rate (FDR) and reasonable sensitivity. However, for achieving a sensitivity of at least 50%, more than 80 samples are required when studying expression levels in a realistic clinical settings such as in cancer research. About half of the methods showed severe excess of false discoveries, making these methods unreliable for differential expression analysis and jeopardizing reproducible science. The detailed results of our study can be consulted through a user-friendly web application, http://statapps.ugent.be/tools/AppDGE/

genomics