bioRxiv ScienceSearch

Biology subjects

Duarte, T.

Publications and source records attributed to Duarte, T..

4 recordsLinked to original sources

A complete nanonpore-only assembly of an XDR Mycobacterium tuberculosis Beijing lineage strain identifies novel genetic variation in repetitive PE/PPE gene regions

A better understanding of the genomic changes that facilitate the emergence and spread of drug resistant M. tuberculosis strains is required. Short-read sequencing methods have limited capacity to identify long, repetitive genomic regions and gene duplications. We sequenced an extensively drug resistant (XDR) Beijing sub-lineage 2.2.1.1 \"epidemic strain\" from the Western Province of Papua New Guinea using long-read sequencing (Oxford Nanopore MinION(R)). With up to 274 fold coverage from a single flow-cell, we assembled a 4404947bp circular genome containing 3670 coding sequences that include the highly repetitive PE/PPE genes. Comparison with Illumina reads indicated a base-level accuracy of 99.95%. Mutations known to confer drug resistance to first and second line drugs were identified and concurred with phenotypic resistance assays. We identified mutations in efflux pump genes (Rv0194), transporters (secA1, glnQ, uspA), cell wall biosynthesis genes (pdk, mmpL, fadD) and virulence genes (mce-gene family, mycp1) that may contribute to the drug resistance phenotype and successful transmission of this strain. Using the newly assembled genome as reference to map raw Illumina reads from representative M. tuberculosis lineages, we detect large insertions relative to the reference genome. We provide a fully annotated genome of a transmissible XDR M. tuberculosis strain from Papua New Guinea using Oxford Nanopore MinION sequencing and provide insight into genomic mechanisms of resistance and virulence.\n\nData SummaryO_LISample Illumina and MinION sequencing reads generated and analyzed are available in NCBI under project accession number PRJNA386696 (https://www.ncbi.nlm.nih.gov/sra/?term=PRJNA386696)\nC_LIO_LIThe assembled complete genome and its annotations are available in NCBI under accession number CP022704.1 (https://www.ncbi.nlm.nih.gov/sra/?term=CP022704.1)\nC_LI\n\nImpact statementWe recently characterized a Modern Beijing lineage strain responsible for the drug resistance outbreaks in the Western province, Papua New Guinea. With some of the genomic markers responsible for its drug resistance and transmissibility are known, there is need to elucidate all molecular mechanisms that account for the resistance phenotype, virulence and transmission. Whole genome sequencing using short reads has widely been utilized to study MTB genome but it does not generally capture long repetitive regions as variants in these regions are eliminated using analysis. Illumina instruments are known to have a GC bias so that regions with high GC or AT rich are under sampled and this effect is exacerbated in MTB, which has approximately 65% GC content. In this study, we utilized Oxford Nanopore Technologies (ONT) MinION sequencing to assemble a high-quality complete genome of an extensively drug resistant strain of a modern Beijing lineage. We were able to able to assemble all PE/PPE (proline-glutamate/proline-proline-glutamate) gene families that have high GC content and repetitive in nature. We show the genomic utility of ONT in offering a more comprehensive understanding of genetic mechanisms that contribute to resistance, virulence and transmission. This is important for settings up predictive analytics platforms and services to support diagnostics and treatment.

genomics

GtTR: Bayesian estimation of absolute tandem repeat copy number using sequence capture and high throughput sequencing

BackgroundTandem repeats comprise significant proportion of the human genome including coding and regulatory regions. They are highly prone to repeat number variation and nucleotide mutation due to their repetitive and unstable nature, making them a major source of genomic variation between individuals. Despite recent advances in high throughput sequencing, analysis of tandem repeats in the context of complex diseases is still hindered by technical limitations.\n\nMethodsWe report a novel targeted sequencing approach, which allows simultaneous analysis of hundreds of repeats. We developed a Bayesian algorithm, namely - GtTR - which combines information from a reference long-read dataset with a short read counting approach to genotype tandem repeats at population scale. PCR sizing analysis was used for validation.\n\nResultsWe used a PacBio long-read sequenced sample to generate a reference tandem repeat genotype dataset with on average 13% absolute deviation from PCR sizing results. Using this reference dataset GtTR generated estimates of VNTR copy number with accuracy within 95% high posterior density (HPD) intervals of 68% and 83% for capture sequence data and 200X WGS data respectively, improving to 87% and 94% with use of a PCR reference. We show that the genotype resolution increases as a function of depth, such that the median 95% HPD interval lies within 25%, 14%, 12% and 8% of the its midpoint copy number value for 30X, 200X WGS, 395X and 800X capture sequence data respectively. We validated nine targets by PCR sizing analysis and genotype estimates from sequencing results correlated well with PCR results.\n\nConclusionsThe novel genotyping approach described here presents a new cost-effective method to explore previously unrecognized class of repeat variation in GWAS studies of complex diseases at the population level. Further improvements in accuracy can be obtained by improving accuracy of the reference dataset.

genomics

Chiron: Translating nanopore raw signal directly into nucleotide sequence using deep learning

Sequencing by translocating DNA fragments through an array of nanopores is a rapidly maturing technology which offers faster and cheaper sequencing than other approaches. However, accurately deciphering the DNA sequence from the noisy and complex electrical signal is challenging. Here, we report Chiron, the first deep learning model to achieve end-to-end basecalling: directly translating the raw signal to DNA sequence without the error-prone segmentation step. Trained with only a small set of 4000 reads, we show that our model provides state-of-the-art basecalling accuracy even on previously unseen species. Chiron achieves basecalling speeds of over 2000 bases per second using desktop computer graphics processing units.

bioinformatics

npInv: accurate detection and genotyping of inversions mediated by non-allelic homologous recombination using long read sub-alignment

Detection of genomic inversions remains challenging. Many existing methods primarily target inversions with a non repetitive breakpoint, leaving inverted repeat (IR) mediated non-allelic homologous recombination (NAHR) inversions largely unexplored. We present npInv, a novel tool specifically for detecting and genotyping NAHR inversion using long read sub-alignment of long read sequencing data. We use npInv to generate a whole-genome inversion map for NA12878 consisting of 30 NAHR inversions (of which 15 are novel), including all previously known NAHR mediated inversions in NA12878 with flanking IR less than 7kb. Our genotyping accuracy on this dataset was 94%. We used PCR to confirm presence of two of these novel NAHR inversions. We show that there is a near linear relationship between the length of flanking IR and the size of the NAHR inversion.

bioinformatics