bioRxiv ScienceSearch

Biology subjects

Gingeras, T. R.

Publications and source records attributed to Gingeras, T. R..

3 recordsLinked to original sources

The fractured landscape of RNA-seq alignment: The default in our STARs

Many tools are available for RNA-seq alignment and expression quantification, with comparative value being hard to establish. Benchmarking assessments often highlight methods good performance, but are focused on either model data or fail to explain variation in performance. This leaves us to ask, what is the most meaningful way to assess different alignment choices? And importantly, where is there room for progress? In this work, we explore the answers to these two questions by performing an exhaustive assessment of the STAR aligner. We assess STARs performance across a range of alignment parameters using common metrics, and then on biologically focused tasks. We find technical metrics such as fraction mapping or expression profile correlation to be uninformative, capturing properties unlikely to have any role in biological discovery. Surprisingly, we find that changes in alignment parameters within a wide range have little impact on both technical and biological performance. Yet, when performance finally does break, it happens in difficult regions, such as X-Y paralogs and MHC genes. We believe improved reporting by developers will help establish where results are likely to be robust or fragile, providing a better baseline to establish where methodological progress can still occur.

bioinformatics

Conserved noncoding transcription and core promoter regulatory code in early Drosophila development

Multicellular development is largely determined by transcriptional regulatory programs that orchestrate the expression of thousands of protein-coding and noncoding genes. To decipher the genomic regulatory code that specifies these programs, and to investigate globally the developmental relevance of noncoding transcription, we profiled genome-wide promoter activity throughout embryonic development in 5 Drosophila species. We show that core promoters, generally not thought to play a significant regulatory role, in fact impart broad restrictions on the developmental timing of gene expression on a genome-wide scale. We propose a hierarchical model of transcriptional regulation during development in which core promoters define broad windows of opportunity for expression, by defining a limited range of transcription factors from which they are able to receive regulatory inputs. This two-tiered mechanism globally orchestrates developmental gene expression, including noncoding transcription on a scale that defies our current understanding of ontogenesis. Indeed, noncoding transcripts are far more prevalent than ever reported before, with [~]4,000 long noncoding RNAs expressed during embryogenesis. Over 1,500 are functionally conserved throughout the melanogaster subgroup, and hundreds are under strong purifying selection. Overall, this work introduces a hierarchical model for the developmental regulation of transcription, and reveals the central role of noncoding transcription in animal development.

genetics

High-throughput annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS)

Accurate annotations of genes and their transcripts is a foundation of genomics, but no annotation technique presently combines throughput and accuracy. As a result, reference gene collections remain incomplete: many gene models are fragmentary, while thousands more remain uncatalogued-particularly for long noncoding RNAs (lncRNAs). To accelerate lncRNA annotation, the GENCODE consortium has developed RNA Capture Long Seq (CLS), combining targeted RNA capture with third-generation long-read sequencing. We present an experimental re-annotation of the GENCODE intergenic lncRNA population in matched human and mouse tissues, resulting in novel transcript models for 3574 / 561 gene loci, respectively. CLS approximately doubles the annotated complexity of targeted loci, outperforming existing short-read techniques. Full-length transcript models produced by CLS enable us to definitively characterize the genomic features of lncRNAs, including promoter- and gene-structure, and protein-coding potential. Thus CLS removes a longstanding bottleneck of transcriptome annotation, generating manual-quality full-length transcript models at high-throughput scales.\n\nAbbreviations

genomics