bioRxiv Science⌕ Search

Biology subjects

van der Toorn, W.

Publications and source records attributed to van der Toorn, W..

2 recordsLinked to original sources

Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing

Nanopore direct RNA sequencing (dRNA-seq) enables unique insights into (epi-)transcriptomics. However, applications are currently limited by the lack of accurate and cost-effective sample multiplexing. We introduce WarpDemuX, an ultra-fast and highly accurate adapter-barcoding and demultiplexing approach. WarpDemuX enhances speed and accuracy by fast processing of the raw nanopore signal, use of a lightweight machine-learning algorithm and design of optimized barcode sets. We demonstrate its utility by performing a rapid phenotypic profiling of different SARS-CoV-2 viruses, crucial for pandemic prevention and response, through multiplexed sequencing of longitudinal samples on a single flowcell. This identifies systematic differences in transcript abundance and poly(A) tail lengths during infection. Additionally, integrating WarpDemuX into sequencing control software enables real-time enrichment of target molecules through barcode-specific adaptive sampling, which we demonstrate by enriching low abundance viral RNA. In summary, WarpDemuX is a broadly applicable, high-performance, and economical multiplexing solution for nanopore dRNA-seq, facilitating advanced (epi-)transcriptomic research.

bioinformatics↗

Sequencing accuracy and systematic errors in nanopore direct RNA sequencing

Direct RNA sequencing (dRNA-seq) on the Oxford Nanopore Technologies (ONT) platforms can produce reads covering up to full-length gene transcripts while containing decipherable information about RNA base modifications and poly-A tail lengths. Although many published studies have been exploring and expanding the potential of dRNA-seq, the sequencing accuracy and error patterns remain understudied. We present the first comprehensive evaluation of accuracy and systematic errors in dRNA-seq data from diverse species, as well as synthetic RNA. Deletions significantly outnumbered mismatches/insertions, while the median read accuracy exhibited species-level variation. In addition to homopolymer errors, we observed systematic biases across nucleotides and heteropolymeric motifs in all species. In general, cytosine/uracil-rich regions were more likely to be erroneous than guanines/adenines. Moreover, the systematic errors were strongly dependent on local sequence contexts. By examining raw signal data, we identified underlying signal-level features potentially associated with the error patterns. While read quality scores approximated error rates at base and read levels, failure to detect DNA adapters may lead to data loss. By comparing distinct basecallers, we reason that some sequencing errors are attributable to signal insufficiency rather than algorithmic (base-calling) artefacts. Lastly, we discuss the implications of such error patterns for downstream applications of dRNA-seq data.

bioinformatics↗