bioRxiv Science⌕ Search

Biology subjects

Furumo, Q.

Publications and source records attributed to Furumo, Q..

2 recordsLinked to original sources

Analysis of 3'-seq data from multiple E. coli studies identifies diverging results sets and raw data characteristics despite similar collection conditions

3-prime end sequencing (3-seq) is a high-throughput sequencing technique that is used to specifically quantify the changes in 3-end formation of transcripts in bacterial cells, which is increasingly being utilized to address fundamental questions regarding transcription termination and pausing across a range of different bacterial species. However, the growing number of 3-seq studies is accompanied by an increase in study-specific 3-seq data analysis approaches. Thus, differences in a number of factors including: experimental design, data collection approaches, analysis methodologies, and interpretation decisions, make it challenging to confidently compare results derived from different studies, even those that were performed on the same organism. To assess the potential severity of these discrepancies, we used PIPETS, a statistically robust and genome-annotation agnostic 3-seq analysis package, to study Escherichia coli 3-seq data sets from three different groups collected under similar conditions. By using a consistent analysis and results interpretation approach, we identified large disparities in the characteristics of the raw 3-seq data between each of the studies, despite all three studies using the same strain and very similar reported experimental conditions. Additionally, we found strand-specific inconsistencies, with some data sets having reference strand 3-seq read coverage distributions that differed greatly from the complement strand within the same replicate. Finally, when the 3-seq distribution profiles of the three E. coli studies are compared to studies from four additional bacteria, we identified 3-seq results clustering patterns that are not explained by phylogenetic similarity between organisms. With the large differences seen between data sets from the same organism as well as the inconsistencies seen between replicates from the same data sets, we urge the field to reconsider the assumptions around 3-seq data homogeneity and move towards consistent analysis approaches, and cautious interpretation of the data. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/658996v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@ca8c67org.highwire.dtl.DTLVardef@1c7c7d4org.highwire.dtl.DTLVardef@11047cdorg.highwire.dtl.DTLVardef@1da1791_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

PIPETS: A statistically robust, gene-annotation agnostic analysis method to study bacterial termination using 3'-end sequencing.

BackgroundOver the last decade the drop in short-read sequencing costs has allowed experimental techniques utilizing sequencing to address specific biological questions to proliferate, oftentimes outpacing standardized or effiective analysis approaches for the data generated. There are growing amounts of bacterial 3-end sequencing data, yet there is currently no commonly accepted analysis methodology for this datatype. Most data analysis approaches are somewhat ad hoc and, despite the presence of substantial signal within annotated genes, focus on genomic regions outside the annotated genes (e.g. 3 or 5 UTRs). Furthermore, the lack of consistent systematic analysis approaches, as well as the absence of genome-wide ground truth data, make it impossible to compare conclusions generated by diffierent labs, using diffierent organisms. ResultsWe present PIPETS, (Poisson Identification of PEaks from Term-Seq data), an R package available on Bioconductor that provides a novel analysis method for 3-end sequencing data. PIPETS is a statistically informed, gene-annotation agnostic methodology. Across two diffierent datasets from two diffierent organisms, PIPETS identified significant 3-end termination signal across a wider range of annotated genomic contexts than existing analysis approaches, suggesting that existing approaches may miss biologically relevant signal. Furthermore, assessment of the previously called 3-end positions not captured by PIPETS showed that they were uniformly very low coverage. ConclusionsPIPETS provides a broadly applicable platform to explore and analyze 3-end sequencing data sets from across diffierent organisms. It requires only the 3-end sequencing data, and is broadly accessible to non-expert users.

bioinformatics↗