bioRxiv · 10.1101/105064
High-throughput annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS)
Abstract
Accurate annotations of genes and their transcripts is a foundation of genomics, but no annotation technique presently combines throughput and accuracy. As a result, reference gene collections remain incomplete: many gene models are fragmentary, while thousands more remain uncatalogued-particularly for long noncoding RNAs (lncRNAs). To accelerate lncRNA annotation, the GENCODE consortium has developed RNA Capture Long Seq (CLS), combining targeted RNA capture with third-generation long-read sequencing. We present an experimental re-annotation of the GENCODE intergenic lncRNA population in matched human and mouse tissues, resulting in novel transcript models for 3574 / 561 gene loci, respectively. CLS approximately doubles the annotated complexity of targeted loci, outperforming existing short-read techniques. Full-length transcript models produced by CLS enable us to definitively characterize the genomic features of lncRNAs, including promoter- and gene-structure, and protein-coding potential. Thus CLS removes a longstanding bottleneck of transcriptome annotation, generating manual-quality full-length transcript models at high-throughput scales.\n\nAbbreviations
Source connections
Explore related subjects
Keep this discovery
Lagarde, J., Uszczynska-Ratajczak, B., Carbonell, S., Davis, C., Gingeras, T. R., Frankish, A., Harrow, J., Guigo, R., Johnson, R.. 2017-02-01. High-throughput annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS). https://doi.org/10.1101/105064
Cite the original work for its findings. Save a collection to share your selection of sources.