bioRxiv ScienceSearch

Biology subjects

Mason, C.

Publications and source records attributed to Mason, C..

5 recordsLinked to original sources

What are the most influencing factors in reconstructing a reliable transcriptome assembly?

Reconstructing the genome and transcriptome for a new or extant species are essential steps in expanding our understanding of the organisms active RNA landscape and gene regulatory dynamics, as well as for developing therapeutic targets to fight disease. The advancement of sequencing technologies has paved the way to generate high-quality draft transcriptomes. With many possible approaches available to accomplish this task, there is a need for a closer investigation of the factors that influence the quality of the results. We carried out an extensive survey of variety of elements that are important in transcriptome assembly. We utilized the human RNA-Seq data from the Sequencing Quality Control Consortium (SEQC) as a well-characterized and comprehensive resource with an available, well-studied human reference genome. Our results indicate that the quality of the library construction significantly impacts the quality of the assembly. Higher coverage of the genome is not as important as the quality of the input RNA-Seq data. Thus, once a certain coverage is attained, the quality of the assembly is mainly dependent on the base-calling accuracy of the input sequencing reads; and it is important to avoid saturating the assembler with extra coverage.

bioinformatics

Minerva: An Alignment and Reference Free Approach to Deconvolve Linked-Reads for Metagenomics

Emerging Linked-Read technologies (aka Read-Cloud or barcoded short-reads) have revived interest in standard short-read technology as a viable way to understand large-scale structure in genomes and metagenomes. Linked-Read technologies, such as the 10X Chromium system, use a microfluidic system and a set of specially designed 3 barcodes (aka UIDs) to tag short DNA reads which were originally sourced from the same long fragment of DNA; subsequently, these specially barcoded reads are sequenced on standard short read platforms. This approach results in interesting compromises. Each long fragment of DNA is covered only sparsely by short reads, no information about the relative ordering of reads from the same fragment is preserved, and typically each 3 barcode matches reads from 2-20 long fragments of DNA. However, compared to long read platforms like those produced by Pacific Biosciences and Oxford Nanopore the cost per base to sequence is far lower, far less input DNA is required, and the per base error rate is that of Illumina short-reads.\n\nThe use of Linked-Reads presents a new set of algorithmic challenges. In this paper, we formally describe one particular issue common to all applications of Linked-Read technology: the deconvolution of reads with a single 3 barcode into clusters that correspond to a single long fragment of DNA. We introduce Minerva, A graph-based algorithm that approximately solves the barcode deconvolution problem for metagenomic data (where reference genomes may be incomplete or unavailable). Additionally, we develop two demonstrations where the deconvolution of barcoded reads improves downstream results: improving the specificity of taxonomic assignments, and by improving clustering of related sequences. To the best of our knowledge, we are the first to address the problem of barcode deconvolution in metagenomics.

bioinformatics

Comprehensive Benchmarking and Ensemble Approaches for Metagenomic Classifiers

BackgroundOne of the main challenges in metagenomics is the identification of microorganisms in clinical and environmental samples. While an extensive and heterogeneous set of computational tools is available to classify microorganisms using whole genome shotgun sequencing data, comprehensive comparisons of these methods are limited. In this study, we use the largest (n=35) to date set of laboratory-generated and simulated controls across 846 species to evaluate the performance of eleven metagenomics classifiers. We also assess the effects of filtering and combining tools to reduce the number of false positives.\n\nResultsTools were characterized on the basis of their ability to (1) identify taxa at the genus, species, and strain levels, (2) quantify relative abundance measures of taxa, and (3) classify individual reads to the species level. Strikingly, the number of species identified by the eleven tools can differ by over three orders of magnitude on the same datasets. However, various strategies can ameliorate taxonomic misclassification, including abundance filtering, ensemble approaches, and tool intersection. Indeed, leveraging tools with different heuristics is beneficial for improved precision. Nevertheless, these strategies were often insufficient to completely eliminate false positives from environmental samples, which are especially important where they concern medically relevant species and where customized tools may be required.\n\nConclusionsThe results of this study provide positive controls, titrated standards, and a guide for selecting tools for metagenomic analyses by comparing ranges of precision and recall. We show that proper experimental design and analysis parameters, including depth of sequencing, choice of classifier or classifiers, database size, and filtering, can reduce false positives, provide greater resolution of species in complex metagenomic samples, and improve the interpretation of results.

genomics

Differing strategies used by motor neurons and glia to achieve robust development of an adult neuropil in Drosophila

In both vertebrates and invertebrates, neurons and glia are generated in a stereotyped order from dedicated progenitors called neural stem cells, but the purpose of invariant lineages is not understood. Here we show that three of the stem cells that produce leg motor neurons in Drosophila also generate a specialized subset of glia, the neuropil glia, which wrap and send processes into the neuropil where motor neuron dendrites arborize. The development of the neuropil glia and leg motor neurons is highly coordinated. However, although individual motor neurons have a stereotyped birth order and transcription factor code, both the number and individual morphologies of the glia born from these lineages are highly plastic, even though the final structure they contribute to is highly stereotyped. We suggest that the shared lineages of these two cell types facilitates the assembly of complex neural circuits, and that the two different birth order strategies - hardwired for motor neurons and flexible for glia - are important for robust nervous system development and homeostasis.

neuroscience

Human PGBD5 DNA transposase promotes site-specific oncogenic mutations in rhabdoid tumors

Genomic rearrangements are a hallmark of childhood solid tumors, but their mutational causes remain poorly understood. Here, we identify the piggyBac transposable element derived 5 (PGBD5) gene as an enzymatically active human DNA transposase expressed in the majority of rhabdoid tumors, a lethal childhood cancer. Using assembly-based whole-genome DNA sequencing, we observed previously unknown somatic genomic rearrangements in primary human rhabdoid tumors. These rearrangements were characterized by deletions and inversions involving PGBD5-specific signal (PSS) sequences at their breakpoints, with some recurrently targeting tumor suppressor genes, leading to their inactivation. PGBD5 was found to be physically associated with human genomic PSS sequences that were also sufficient to mediate PGBD5-induced DNA rearrangements in rhabdoid tumor cells. We found that ectopic expression of PGBD5 in primary immortalized human cells was sufficient to promote penetrant cell transformation in vitro and in immunodeficient mice in vivo. This activity required specific catalytic residues in the PGBD5 transposase domain, as well as end-joining DNA repair, and induced distinct structural rearrangements, involving PSS-associated breakpoints, similar to those found in primary human rhabdoid tumors. This defines PGBD5 as an oncogenic mutator and provides a plausible mechanism for site-specific DNA rearrangements in childhood and adult solid tumors.

cancer biology