bioRxiv ScienceSearch

Biology subjects

Egan, R.

Publications and source records attributed to Egan, R..

5 recordsLinked to original sources

MiniScrub: de novo long read scrubbing using approximate alignment and deep learning

Long read sequencing technologies such as Oxford Nanopore can greatly de-crease the complexity of de novo genome assembly and large structural variation iden-tification. Currently Nanopore reads have high error rates, and the errors often cluster into low-quality segments within the reads. Many methods for resolving these errors require access to reference genomes, high-fidelity short reads, or reference genomes, which are often not available. De novo error correction modules are available, often as part of assembly tools, but large-scale errors still remain in resulting assemblies, motivating further innovation in this area. We developed a novel Convolutional Neu-ral Network (CNN) based method, called MiniScrub, for de novo identification and subsequent \"scrubbing\" (removal) of low-quality Nanopore read segments. MiniScrub first generates read-to-read alignments by MiniMap, then encodes the alignments into images, and finally builds CNN models to predict low-quality segments that could be scrubbed based on a customized quality cutoff. Applying MiniScrub to real world con-trol datasets under several different parameters, we show that it robustly improves read quality. Compared to raw reads, de novo genome assembly with scrubbed reads pro-duces many fewer mis-assemblies and large indel errors. We propose MiniScrub as a tool for preprocessing Nanopore reads for downstream analyses. MiniScrub is open-source software and is available at https://bitbucket.org/berkeleylab/jgi-miniscrub

bioinformatics

Head-to-Head Comparison of Three Methods of Quantifying Competitive Fitness in C. elegans

Organismal fitness is relevant in many contexts in biology. The most meaningful experimental measure of fitness is competitive fitness, when two or more entities (e.g., genotypes) are allowed to compete directly. In theory, competitive fitness is simple to measure: an experimental population is initiated with the different types in known proportions and allowed to evolve under experimental conditions to a predefined endpoint. In practice, there are several obstacles to obtaining robust estimates of competitive fitness in multicellular organisms, the most pervasive of which is simply the time it takes to count many individuals of different types from many replicate populations. Methods by which counting can be automated in high throughput are desirable, but for automated methods to be useful, the bias and technical variance associated with the method must be (a) known, and (b) sufficiently small relative to other sources of bias and variance to make the effort worthwhile.\n\nThe nematode Caenorhabditis elegans is an important model organism, and the fitness effects of genotype and environmental conditions are often of interest. We report a comparison of three experimental methods of quantifying competitive fitness, in which wild-type strains are competed against GFP-marked competitors under standard laboratory conditions. Population samples were split into three replicates and counted (1) \"by eye\" from a saved image, (2) from the same image using CellProfiler image analysis software, and (3) with a large particle flow cytometer (a \"worm sorter\"). From 720 replicate samples, neither the frequency of wild-type worms nor the among-sample variance differed significantly between the three methods. CellProfiler and the worm sorter provide at least a tenfold increase in sample handling speed with little (if any) bias or increase in variance.

evolutionary biology

Genome-wide prediction of synthetic rescue mediators of resistance to targeted and immunotherapy

Most patients with advanced cancer eventually acquire resistance to targeted therapies, spurring extensive efforts to identify molecular events mediating therapy resistance. Many of these events involve synthetic rescue (SR) interactions, where the reduction in cancer cell viability caused by targeted gene inactivation is rescued by an adaptive alteration of another gene (the rescuer). Here we perform a genome-wide prediction of SR rescuer genes by analyzing tumor transcriptomics and survival data of 10,000 TCGA cancer patients. Predicted SR interactions are validated in new experimental screens. We show that SR interactions can successfully predict cancer patients response and emerging resistance. Inhibiting predicted rescuer genes sensitizes resistant cancer cells to therapies synergistically, providing initial leads for developing combinatorial approaches to overcome resistance proactively. Finally, we show that the SR analysis of melanoma patients successfully identifies known mediators of resistance to immunotherapy and predicts novel rescuers.

cancer biology

Critical Assessment of Metagenome Interpretation - a benchmark of computational metagenomics software

In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.

bioinformatics

De novo Identification of DNA Modifications Enabled by Genome-Guided Nanopore Signal Processing

Advances in nanopore sequencing technology have enabled investigation of the full catalogue of covalent DNA modifications. We present the first algorithm for the identification of modified nucleotides without the need for prior training data along with the open source software implementation, nanoraw. Nanoraw accurately assigns contiguous raw nanopore signal to genomic positions, enabling novel data visualization, and increasing power and accuracy for the discovery of covalently modified bases in native DNA. Ground truth case studies utilizing synthetically methylated DNA show the capacity to identify three distinct methylation marks, 4mC, 5mC, and 6mA, in seven distinct sequence contexts without any changes to the algorithm. We demonstrate quantitative reproducibility simultaneously identifying 5mC and 6mA in native E. coli across biological replicates processed in different labs. Finally we propose a pipeline for the comprehensive discovery of DNA modifications in any genome without a priori knowledge of their chemical identities.

bioinformatics