bioRxiv ScienceSearch

Biology subjects

W. Richard McCombie

Publications and source records attributed to W. Richard McCombie.

4 recordsLinked to original sources

SiLiCO: A Simulator of Long Read Sequencing in PacBio and Oxford Nanopore

SummaryLong read sequencing platforms, which include the widely used Pacific Biosciences (PacBio) platform and the emerging Oxford Nanopore platform, aim to produce sequence fragments in excess of 15-20 kilobases, and have proved advantageous in the identification of structural variants and easing genome assembly. However, long read sequencing remains relatively expensive and error prone, and failed sequencing runs represent a significant problem for genomics core facilities. To quantitatively assess the underlying mechanics of sequencing failure, it is essential to have highly reproducible and controllable reference data sets to which sequencing results can be compared. Here, we present SiLiCO, the first in silico simulation tool to generate standardized sequencing results from both of the leading long read sequencing platforms.\n\nAvailabilitySiLiCO is an open source package written in Python. It is freely available at https://www.github.com/ethanagbaker/SiLiCO under the GNU GPL 3.0 license.\n\nContact \n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics

C. elegans PVD Neurons: A Platform for Functionally Validating and Characterizing Neuropsychiatric Risk Genes

One of the primary challenges in the field of psychiatric genetics is the lack of an in vivo model system in which to functionally validate candidate neuropsychiatric risk genes (NRGs) in a rapid and cost-effective manner1-3. To overcome this obstacle, we performed a candidate-based RNAi screen in which C. elegans orthologs of human NRGs were assayed for dendritic arborization and cell specification defects using C. elegans PVD neurons. Of 66 NRGs, identified via exome sequencing of autism (ASD)4 or schizophrenia (SCZ)5-9 probands and whose mutations are de novo and predicted to result in a complete or partial loss of protein function, the C. elegans orthologs of 7 NRGs were found to be required for proper neuronal development and represent a variety of functional classes, including transcriptional regulators and chromatin remodelers, molecular chaperones, and cytoskeleton-related proteins. Notably, the positive hit rate, when selectively assaying C. elegans orthologs of ASD and SCZ NRGs, is enriched >14-fold as compared to unbiased RNAi screening10. Furthermore, we find that RNAi phenotypes associated with the depletion of NRG orthologs is recapitulated in genetic mutant animals, and, via genetic interaction studies, we show that the NRG ortholog of ANK2, unc-44, is required for SAX-7/MNR-1/DMA-1 signaling. Collectively, our studies demonstrate that C. elegans PVD neurons are a tractable model in which to discover and dissect the fundamental molecular mechanisms underlying neuropsychiatric disease pathogenesis.

Neuroscience

Third-generation sequencing and the future of genomics

Third-generation long-range DNA sequencing and mapping technologies are creating a renaissance in high-quality genome sequencing. Unlike second-generation sequencing, which produces short reads a few hundred base-pairs long, third-generation single-molecule technologies generate over 10,000 bp reads or map over 100,000 bp molecules. We analyze how increased read lengths can be used to address longstanding problems in de novo genome assembly, structural variation analysis and haplotype phasing.

Bioinformatics

Error correction and assembly complexity of single molecule sequencing reads.

Third generation single molecule sequencing technology is poised to revolutionize genomics by enabling the sequencing of long, individual molecules of DNA and RNA. These technologies now routinely produce reads exceeding 5,000 basepairs, and can achieve reads as long as 50,000 basepairs. Here we evaluate the limits of single molecule sequencing by assessing the impact of long read sequencing in the assembly of the human genome and 25 other important genomes across the tree of life. From this, we develop a new data-driven model using support vector regression that can accurately predict assembly performance. We also present a novel hybrid error correction algorithm for long PacBio sequencing reads that uses pre-assembled Illumina sequences for the error correction. We apply it several prokaryotic and eukaryotic genomes, and show it can achieve near-perfect assemblies of small genomes (< 100Mbp) and substantially improved assemblies of larger ones. All source code and the assembly model are available open-source.

Bioinformatics