bioRxiv Science⌕ Search

Biology subjects

Stanojevic, D.

Publications and source records attributed to Stanojevic, D..

2 recordsLinked to original sources

Rockfish: A Transformer-based Model for Accurate 5-Methylcytosine Prediction from Nanopore Sequencing

DNA methylation plays a crucial role in various biological processes, including cell differentiation, ageing, and cancer development. The most important methylation in mammals is 5-methylcytosine (5mC) which is present in the context of CpG dinucleotides. Sequencing methods such as whole-genome bisulfite sequencing (WGBS) successfully detect 5mC DNA modifications. However, they suffer from the serious drawbacks of short read lengths and might introduce an amplification bias. Here we present Rockfish, a deep learning algorithm that significantly improves read-level 5mC detection by using Nanopore sequencing. Compared to other methods based on Nanopore sequencing, there is an increase in the single-base accuracy and the F1 measure of up to 5% and 12%, respectively. Furthermore, Rockfish shows a high correlation with WGBS and requires lower read depth while being computationally efficient. We deem that Rockfish is broadly applicable to study 5mC methylation in diverse organisms and disease systems to yield biological insights.

bioinformatics↗

A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines

The human genome contains more than 200,000 gene isoforms. However, different isoforms can be highly similar, and with an average length of 1.5kb remain difficult to study with short read sequencing. To systematically evaluate the ability to study the transcriptome at a resolution of individual isoforms we profiled 5 human cell lines with short read cDNA sequencing and Nanopore long read direct RNA, amplification-free direct cDNA, PCR-cDNA sequencing. The long read protocols showed a high level of consistency, with amplification-free RNA and cDNA sequencing being most similar. While short and long reads generated comparable gene expression estimates, they differed substantially for individual isoforms. We find that increased read length improves read-to-transcript assignment, identifies interactions between alternative promoters and splicing, enables the discovery of novel transcripts from repetitive regions, facilitates the quantification of full-length fusion isoforms and enables the simultaneous profiling of m6A RNA modifications when RNA is sequenced directly. Our study demonstrates the advantage of long read RNA sequencing and provides a comprehensive resource that will enable the development and benchmarking of computational methods for profiling complex transcriptional events at isoform-level resolution.

genomics↗