bioRxiv Science⌕ Search

Biology subjects

Brooks, T. G.

Publications and source records attributed to Brooks, T. G..

3 recordsLinked to original sources

BEERS2: RNA-Seq simulation through high fidelity in silico modeling

Simulation of RNA-seq reads is critical in the assessment, comparison, benchmarking, and development of bioinformatics tools. Yet the field of RNA-seq simulators has progressed little in the last decade. To address this need we have developed BEERS2, which combines a flexible and highly configurable design with detailed simulation of the entire library preparation and sequencing pipeline. BEERS2 takes input transcripts (typically fully-length mRNA transcripts with polyA tails) from either customizable input or from CAMPAREE simulated RNA samples. It produces realistic reads of these transcripts as FASTQ, SAM, or BAM formats with the SAM or BAM formats containing the true alignment to the reference genome. It also produces true transcript-level quantification values. BEERS2 combines a flexible and highly configurable design with detailed simulation of the entire library preparation and sequencing pipeline and is designed to include the effects of polyA selection and RiboZero for ribosomal depletion, hexamer priming sequence biases, GC-content biases in PCR amplification, barcode read errors, and errors during PCR amplification. These characteristics combine to make BEERS2 the most complete simulation of RNA-seq to date. Finally, we demonstrate the use of BEERS2 by measuring the effect of several settings on the popular Salmon pseudoalignment algorithm.

bioinformatics↗

Meta-analysis of diurnal transcriptomics reveals strong patterns of concordance and discordance in mouse liver

The accumulation of public transcriptomic timeseries data enables robust meta-analyses that were not possible until recently. To assess the consistency of biological rhythms across studies, 43 public mouse liver tissue timeseries totaling 805 RNA-seq samples were obtained and analyzed. Only the control groups of each study were included, in order to create comparable data. Technical factors in RNA-seq library preparation were the largest contributors to transcriptome-level differences, beyond biological or experiment-specific factors such as lighting conditions. Core clock genes were remarkably consistent in phase across all studies, while phase distributions of other periodic genes were generally less consistent. Overlap of genes identified as rhythmic across studies was generally low, up to around 50% between some of the highest sample count studies. Distributions of phases of significant genes were remarkably inconsistent across studies, but genes consistently identified as rhythmic clustered near ZT0 and ZT12 in acrophase. Data was integrated across studies in a JIVE analysis, which showed that the top two components of joint within-study variation are determined by time of day. A shape-invariant model with random effects was fit to the genes to identify the underlying shape of the rhythms, consistent across all studies. This revealed the extent of asymmetric and multimodal genes.

genomics↗

Comparative evaluation of full-length isoform quantification from RNA-Seq

Full-length isoform quantification from RNA-Seq is a key goal in transcriptomics analyses and has been an area of active development since the beginning. The fundamental difficulty stems from the fact that RNA transcripts are long, while RNA-Seq reads are short. Here we use simulated benchmarking data that reflects many properties of real data, including polymorphisms, intron signal and non-uniform coverage, allowing for systematic comparative analyses of isoform quantification accuracy and its impact on differential expression analysis. Genome, transcriptome and pseudo alignment-based methods are included; and a simple approach is included as a baseline control. Salmon, kallisto, RSEM, and Cufflinks exhibit the highest accuracy on idealized data, while on more realistic data they do not perform dramatically better than the simple approach. We determine the structural parameters with the greatest impact on quantification accuracy to be length and sequence compression complexity and not so much the number of isoforms. The effect of incomplete annotation on performance is also investigated. Overall, the tested methods show sufficient divergence from the truth to suggest that full-length isoform quantification and isoform level DE should still be employed selectively.

bioinformatics↗