bioRxiv ScienceSearch

Biology subjects

Cao, M. D.

Publications and source records attributed to Cao, M. D..

9 recordsLinked to original sources

Octapeptin C4 Induces Less Resistance and Novel Mutations in an Epidemic Carbapenemase-producing Klebsiella pneumoniae ST258 Clinical Isolate Compared to Polymyxins

Polymyxin B and E (colistin) have been pivotal in the treatment of extensively drug-resistant (XDR) Gram-negative bacterial infections, with increasing use over the past decade. Unfortunately, resistance to these antibiotics is rapidly emerging. The structurally-related octapeptin C4 (OctC4) has shown significant potency against XDR bacteria, including against polymyxin-resistant (Pmx-R) strains, but its mode of action remains undefined. We sought to compare and contrast the acquisition of XDR Klebsiella pneumoniae (ST258) resistance in vitro with all three lipopeptides to help elucidate the mode of action of the drugs and potential mechanisms of resistance evolution. Strikingly, 20 days of exposure to the polymyxins resulted in a dramatic (1000-fold) increase in the minimum inhibitory concentration (MIC) for the polymyxins, reflecting the evolution of resistance seen in clinical isolates, whereas for OctC4 only a 4-fold increase was witnessed. There was no cross-resistance observed between the polymyxin - and octapeptin-induced resistant strains. Sequencing revealed previously known gene alterations for polymyxin resistance, including crrB, mgrB, pmrB, phoPQ and yciM, and novel mutations in qseC. In contrast, mutations in mlaDF and pqiB, 1genes related to phospholipid transport, were found in octapeptin-resistant isolates. Mutation effects were validated via complementation assays. These genetic variations were reflected in phenotypic changes to lipid A. Pmx-R isolates increased 4-amino-4-deoxy-arabinose fortification to phosphate groups of lipid A, whereas OctC4 induced strains harbored a higher abundance of hydroxymyristate and palmitoylate. The results reveal a differing mode of action compared to polymyxins which provides hope for future therapeutics to combat the increasingly threat of XDR bacteria.

microbiology

GtTR: Bayesian estimation of absolute tandem repeat copy number using sequence capture and high throughput sequencing

BackgroundTandem repeats comprise significant proportion of the human genome including coding and regulatory regions. They are highly prone to repeat number variation and nucleotide mutation due to their repetitive and unstable nature, making them a major source of genomic variation between individuals. Despite recent advances in high throughput sequencing, analysis of tandem repeats in the context of complex diseases is still hindered by technical limitations.\n\nMethodsWe report a novel targeted sequencing approach, which allows simultaneous analysis of hundreds of repeats. We developed a Bayesian algorithm, namely - GtTR - which combines information from a reference long-read dataset with a short read counting approach to genotype tandem repeats at population scale. PCR sizing analysis was used for validation.\n\nResultsWe used a PacBio long-read sequenced sample to generate a reference tandem repeat genotype dataset with on average 13% absolute deviation from PCR sizing results. Using this reference dataset GtTR generated estimates of VNTR copy number with accuracy within 95% high posterior density (HPD) intervals of 68% and 83% for capture sequence data and 200X WGS data respectively, improving to 87% and 94% with use of a PCR reference. We show that the genotype resolution increases as a function of depth, such that the median 95% HPD interval lies within 25%, 14%, 12% and 8% of the its midpoint copy number value for 30X, 200X WGS, 395X and 800X capture sequence data respectively. We validated nine targets by PCR sizing analysis and genotype estimates from sequencing results correlated well with PCR results.\n\nConclusionsThe novel genotyping approach described here presents a new cost-effective method to explore previously unrecognized class of repeat variation in GWAS studies of complex diseases at the population level. Further improvements in accuracy can be obtained by improving accuracy of the reference dataset.

genomics

Chiron: Translating nanopore raw signal directly into nucleotide sequence using deep learning

Sequencing by translocating DNA fragments through an array of nanopores is a rapidly maturing technology which offers faster and cheaper sequencing than other approaches. However, accurately deciphering the DNA sequence from the noisy and complex electrical signal is challenging. Here, we report Chiron, the first deep learning model to achieve end-to-end basecalling: directly translating the raw signal to DNA sequence without the error-prone segmentation step. Trained with only a small set of 4000 reads, we show that our model provides state-of-the-art basecalling accuracy even on previously unseen species. Chiron achieves basecalling speeds of over 2000 bases per second using desktop computer graphics processing units.

bioinformatics

npInv: accurate detection and genotyping of inversions mediated by non-allelic homologous recombination using long read sub-alignment

Detection of genomic inversions remains challenging. Many existing methods primarily target inversions with a non repetitive breakpoint, leaving inverted repeat (IR) mediated non-allelic homologous recombination (NAHR) inversions largely unexplored. We present npInv, a novel tool specifically for detecting and genotyping NAHR inversion using long read sub-alignment of long read sequencing data. We use npInv to generate a whole-genome inversion map for NA12878 consisting of 30 NAHR inversions (of which 15 are novel), including all previously known NAHR mediated inversions in NA12878 with flanking IR less than 7kb. Our genotyping accuracy on this dataset was 94%. We used PCR to confirm presence of two of these novel NAHR inversions. We show that there is a near linear relationship between the length of flanking IR and the size of the NAHR inversion.

bioinformatics

Multifactorial Chromosomal Variants Regulate Polymyxin Resistance In Extensively Drug-Resistant Klebsiella pneumoniae

Extensively drug-resistant Klebsiella pneumoniae (XDR-KP) infections cause high mortality and are disseminating globally. Identifying the genetic basis underpinning resistance allows for rapid diagnosis and treatment. XDR isolates sourced from Greece and Brazil, including nineteen polymyxin-resistant and five polymyxin-susceptible strains, underwent whole genome sequencing. Approximately 90% of polymyxin resistance was enabled by alterations upstream or within mgrB. The most common mutation identified was an insertion at nucleotide position 75 in mgrB via an ISKpn26-like element in the ST258 lineage and ISKpn13 in one ST11 isolate. Three strains acquired an IS1 element upstream of mgrB and another strain had an ISKpn25 insertion at 133 bp. Other isolates had truncations (C28STOP, Q30STOP) or a missense mutation (D31E) affecting mgrB. Complementation assays revealed all mgrB perturbations contributed to resistance. Missense mutations in phoQ (T281M, G385C) were also found to facilitate resistance. Several variants in phoPQ co-segregating with the ISKpn26-like insertion were identified as potential partial suppressor mutations. Three ST258 samples were found to contain subpopulations with different resistance conferring mutations, including the ISKpn26-like insertion colonising with a novel mutation in pmrB (P158R), both confirmed via complementation assays. We also characterized a new multi-drug resistant Klebsiella quasipneumoniae strain ST2401 which was susceptible to polymyxins. These findings highlight the broad spectrum of chromosomal modifications which can facilitate and regulate resistance against polymyxins in K. pneumoniae.\n\nDATA SUMMARYO_LIWhole genome sequencing of the 24 clinical isolates has been deposited under BioProject PRJNA307517 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA307517).\nC_LI\n\nIMPACT STATEMENTKlebsiella pneumoniae contributes to a high abundance of nosocomial infections and the rapid emergence of antimicrobial resistance hinders treatment. Polymyxins are predominantly utilized to treat multidrug-resistant infections, however, resistance to the polymyxins is arising. This increasing prevalence in polymyxin resistance is evident especially in Greece and Brazil. Identifying the genomic variations conferring resistance in clinical isolates from these regions assists with potentially detecting novel alterations and tracing the spread of particular strains. This study commonly found mutations in the gene mgrB, the negative regulator of PhoPQ, known to cause resistance in KP. In the remaining isolates, missense mutations in phoQ were accountable for resistance. Multiple novel mutations were detected to be segregating with mgrB perturbations. This was either due to a mixed heterogeneous sample of two polymyxin-resistant strains, or because of multiple mutations within the same strain. Of interest was the validation of novel mutations inphoPQ segregating with a previously known ISKpn26-like element in disrupted mgrB isolates. Complementation of these phoPQ mutations revealed a reduction in minimum inhibitory concentrations and suggests the first evidence of partial suppressor mutations in KP. This research builds upon our current understanding of heteroresistance, lineage specific mutations and regulatory variations relating to polymyxin resistance.

microbiology

Simulating The Dynamics Of Targeted Capture Sequencing With CapSim

MotivationTargeted sequencing using capture probes has become increasingly popular in clinical applications due to its scalability and cost-effectiveness. The approach also allows for higher sequencing coverage of the targeted regions resulting in better analysis statistical power. However, because of the dynamics of the hybridisation process, it is difficult to evaluate the efficiency of the probe design prior to the experiments which are time consuming and costly.\n\nResultsWe developed CapSim, a software package for simulation of targeted sequencing. Given a genome sequence and a set of probes, CapSim simulates the fragmentation, the dynamics of probe hybridisation, and the sequencing of the captured fragments on Illumina and PacBio sequencing platforms. The simulated data can be used for evaluating the performance of the analysis pipeline, as well as the efficiency of the probe design. Parameters of the various stages in the sequencing process can also be evaluated in order to optimise the efficacy of the experiments.\n\nAvailabilityCapSim is publicly available under BSD license at https://github.com/mdcao/capsim.

bioinformatics

Real-Time Demultiplexing Nanopore Barcoded Sequencing Data With npBarcode

MotivationThe recently introduced barcoding protocol to Oxford Nanopore sequencing has increased the versatility of the technology. Several bioinformatic tools have been developed to demultiplex the barcoded reads, but none of them support the streaming analysis. This limits the use of pooled sequencing in real-time applications, which is one of the main advantages of the technology.\n\nResultsWe introduced npBarcode, an open source and cross platform tool for barcode demultiplex in streaming fashion. npBarcode can be seamlessly integrated into a streaming analysis pipeline. The tool also provides a friendly graphical user interface through npReader, allowing the real-time visual monitoring of the sequencing progress of barcoded samples. We show that npBarcode achieves comparable accuracies to the other alternatives.\n\nAvailabilitynpBarcode is bundled in Japsa - a Java tools kit for genome analysis, and is freely available at https://github.com/hsnguyen/npBarcode.

bioinformatics

Assembly Of Whole-Chromosome Pseudomolecules For Polyploid Plant Genomes Using Outcrossed Mapping Populations

The assembly of whole-chromosome pseudomolecules for plant genomes remains challenging due to polyploidy and high repeat content. We developed an approach for constructing complete pseudomolecules for polyploid species using genotyping-by-sequencing data from outcrossing mapping populations coupled with high coverage whole genome sequence data of a reference genome. Our approach combines de novo assembly with linkage mapping to arrange scaffolds into pseudomolecules. We show that the method is able to reconstruct simulated chromosomes for both diploid and tetraploid genomes. Comparisons to three existing genetic mapping tools show that our method outperforms the other methods in accuracy on both grouping and ordering, and is robust to the presence of substantial amounts of missing data and genotyping errors. We applied our method to three real datasets including a diploid Ipomoea trifida and two tetraploid potato mapping populations. The linkage maps show significant concordance with the reference chromosomes. We resolved seven assembly errors for the published Ipomoea trifida genome assembly as well as anchored an unplaced scaffold in the published potato genome.

bioinformatics

Ongoing human chromosome end extension driven by a primate ancestral genomic region revealed by analysis of BioNano genomics data

The majority of human chromosome ends remain incompletely assembled due to their highly repetitive structure. In this study, we use BioNano data to anchor and extend chromosome ends from two European trios as well as two unrelated Asian genomes. BioNano assembled chromosome ends are structurally divergent from the reference genome, including both missing sequence (10%) and extensions(22%). These extensions are heritable and in some cases divergent between Asian and European samples. Six ninths of the extension sequence in NA12878 can be confirmed and filled by nanopore data. We identify two sequence families in these sequences which have undergone substantial duplication in multiple primate lineages. We show that these sequence families have arisen from progenitor interstitial sequence on the ancestral primate chromosome 7. Comparison of chromosome end sequences from 15 species revealed that chromosome end missing sequence matches the corresponding phylogenetic relationship and revealed a rate of chromosome extension per chromosome of 0.0020 bp per year in average.

genomics