bioRxiv Science⌕ Search

Biology subjects

Jafarpour, S.

Publications and source records attributed to Jafarpour, S..

4 recordsLinked to original sources

The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding

Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.

bioinformatics↗

Performance Test of the QNome Nanopore Sequencer

Nanopore sequencers have the potential to liberate DNA sequencing from centralized core facilities to distributed analytical nodes. Until now, Oxford Nanopore Technologies (ONT) has been the sole manufacturer of a portable nanopore sequencer, but analogous platforms are in production. Nanopore sequencers from Qitan Technology (QT) are widely used in China but have been unavailable outside that nation and lack independent performance testing. Enabled by early access to QTs least expensive sequencer and flow cell, the QNome-3841 and QCell-384, we tested whether they could generate accurate DNA barcodes cost-effectively. In several tests involving amplicon pools from 95 to 9,120 specimens, QT recovered valid DNA barcodes from nearly as many specimens (98%) as ONT. QT sequences had slightly lower fidelity than their ONT counterparts and QT frequently failed to resolve the correct length of G/C homopolymers. However, barcode sequences from the two platforms were nearly indistinguishable after bioinformatic treatment. QTs wash kit performed well, enabling a QCell to sequence eight amplicon pools with zero carryover between runs and minimal degradation of the flow cell. Its ultra-fast protocol allowed library preparation in a single step that could be completed in 15 minutes, but this came at the cost of lower quality data. Once widely available, QT devices will be well-suited for supporting DNA barcode analysis.

genomics↗

Incidence and attributes of chimeric COI and 18S sequences derived from nematode-infected arthropods

High-throughput sequencing is speeding the discovery of new species and making it possible to quantify the species richness of bulk samples. However, chimeric sequences can lead to both type I errors (true species mistakenly rejected) and type II errors (chimeras incorrectly viewed as valid species). This study employed nanopore sequencing to examine the incidence and nature of chimeric sequences recovered from 531 arthropod specimens infected with nematodes. Specifically, it examined chimera formation in two gene regions: the 658 bp segment of mitochondrial COI employed as the barcode region for the animal kingdom, and a 900 bp segment of nuclear 18S often employed to discriminate lineages of nematodes. This work revealed chimeras for both gene regions despite the deep sequence divergences between members of these two phyla. However, the incidence of chimeric molecules was higher for 18S than COI, with chimera formation correlating with local variation in DNA stability and conserved regions between parent sequences. Aside from demonstrating that chimeras can arise from distantly related organisms, this study provides insights into the mechanisms underlying their formation.

molecular biology↗

Barcode 100K Specimens: In a Single Nanopore Run

It is a global priority to better manage the biosphere, but action needs to be informed by monitoring shifts in the abundance and distribution of species across the domains of life. The acquisition of such information is currently constrained by the limited knowledge of biodiversity. Among the 20 million or more species of eukaryotes, just a tenth have scientific names. DNA barcoding can speed the registration of unknown animal species, the most diverse kingdom of eukaryotes, as the BIN system automates their recognition. However, inexpensive analytical protocols are critical as the census of all animal species will require processing a billion or more specimens. Barcoding involves DNA extraction followed by PCR and sequencing with the last step dominating costs until 2017. By recovering barcodes from highly multiplexed samples, the Sequel platforms from Pacific BioSciences slashed costs by 90%, but these instruments are only deployed in core facilities because of their expense. Sequencers from Oxford Nanopore Technologies provide an escape from high capital and service costs, but their low sequence fidelity has, until now, kept analytical cost above Sequel. However, the improved performance of its latest flow cells (R10.4.1) might erase this differential. This study demonstrates that a regular MinION flow cell can characterize an amplicon pool derived from 100,000 specimens while a Flongle flow cell can process one derived from several thousand. At $0.01 per specimen, DNA sequencing is now the least expensive step in the barcode workflow. By coupling simplified protocols for DNA extraction with ultra-low volume PCRs, it will be possible to move from specimen to DNA barcode for $0.10, a price point that will enable the census of all species within two decades.

molecular biology↗