bioRxiv · 10.1101/2021.12.08.471868
Improved Transcriptome Assembly Using a Hybrid of Long and Short Reads with StringTie
Abstract
Short-read RNA sequencing and long-read RNA sequencing each have their strengths and weaknesses for transcriptome assembly. While short reads are highly accurate, they are unable to span multiple exons. Long-read technology can capture full-length transcripts, but its high error rate often leads to mis-identified splice sites, and its low throughput makes quantification difficult. Here we present a new release of StringTie that performs hybrid-read assembly. By taking advantage of the strengths of both long and short reads, hybrid-read assembly with StringTie is more accurate than long-read only or short-read only assembly, and on some datasets it can more than double the number of correctly assembled transcripts, while obtaining substantially higher precision than the long-read data assembly alone. Here we demonstrate the improved accuracy on simulated data and real data from Arabidopsis thaliana, Mus musculus, and human. We also show that hybrid-read assembly is more accurate than correcting long reads prior to assembly while also being substantially faster. StringTie is freely available as open source software at https://github.com/gpertea/stringtie.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shumate, A., Wong, B., Pertea, G., Pertea, M.. 2021-12-10. Improved Transcriptome Assembly Using a Hybrid of Long and Short Reads with StringTie. https://doi.org/10.1101/2021.12.08.471868
Cite the original work for its findings. Save a collection to share your selection of sources.