bioRxiv ScienceSearch

Biology subjects

DiCarlo, J.

Publications and source records attributed to DiCarlo, J..

2 recordsLinked to original sources

smCounter2: an accurate low-frequency variant caller for targeted sequencing data with unique molecular identifiers

MotivationLow-frequency DNA mutations are often confounded with technical artifacts from sample preparation and sequencing. With unique molecular identifiers (UMIs), most of the sequencing errors can be corrected. However, errors before UMI tagging, such as DNA polymerase errors during end-repair and the first PCR cycle, cannot be corrected with single-strand UMIs and impose fundamental limits to UMI-based variant calling.\n\nResultsWe developed smCounter2, a UMI-based variant caller for targeted sequencing data and an upgrade from the current version of smCounter. Compared to smCounter, smCounter2 features lower detection limit at 0.5%, better overall accuracy (particularly in non-coding regions), a consistent threshold that can be applied to both deep and shallow sequencing runs, and easier use via a Docker image and code for read pre-processing. We benchmarked smCounter2 against several state-of-the-art UMI-based variant calling methods using multiple datasets and demonstrated smCounter2s superior performance in detecting somatic variants. At the core of smCounter2 is a statistical test to determine whether the allele frequency of the putative variant is significantly above the background error rate, which was carefully modeled using an independent dataset. The improved accuracy in non-coding regions was mainly achieved using novel repetitive region filters that were specifically designed for UMI data.\n\nAvailabilityThe entire pipeline is available at https://github.com/qiaseq/qiaseq-dna under MIT license.

bioinformatics

High-throughput creation and functional profiling of eukaryotic DNA sequence variant libraries using CRISPR/Cas9

Construction of genetic variant libraries with phenotypic measurement is central to advancing todays functional genomics, and remains a grand challenge. Here, we introduce a Cas9-based approach for generating pools of mutants with defined genetic alterations (deletions, substitutions and insertions), along with methods for tracking their fitness en masse. We demonstrate the utility of our approach in performing focused analysis of hundreds of mutants of a single protein and in investigating the biological function of an entire family of poorly characterized genetic elements. Our platform allows fundamental biology questions to be investigated in a quick, easy and affordable manner.

bioengineering