bioRxiv Science⌕ Search

Biology subjects

Loubser, J.

Publications and source records attributed to Loubser, J..

2 recordsLinked to original sources

MTBseq-nf: Enabling Scalable Tuberculosis Genomics "Big Data" Analysis through a User-Friendly Nextflow Wrapper for MTBseq pipeline

The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tu- berculosis research using whole-genome sequencing data. It was the first publicly available pipeline on GitHub to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality control through mapping, variant calling for lineage classifica- tion, drug resistance prediction, and phylogenetic inference. However, the pipelines architecture is not optimal for analyses on high-performance computing or cloud computing environments, which often involve large datasets. To optimize the pipeline, we created MTBseq-nf, a Nextflow wrapper which offers shorter execution times through parallelization along with multiple other key improvements. The MTBseq-nf wrapper, as opposed to the linear, batched analysis of samples in the TBfull step of MTBseq pipeline, can execute multiple instances of the same step in parallel and therefore makes full use of the provided computational resources. For evaluation of scalability and reproducibility, we used 90 M. tuberculosis genomes (European Nucleotide Archive - ENA- accession PRJEB7727) for the benchmarking analysis on a dedicated computational server. In our experiments the execution time of MTBseq-nf parallel analysis mode is at least twice as fast as the standard MTBseq pipeline for more than 20 samples. Furthermore, the MTBseq-nf wrapper facilitates reproducibility using the nf-core, bioconda, and biocontainers projects for platform independence. The proposed MTBseq-nf wrapper pipeline is, user-friendly, optimized for hardware efficiency, scalable for larger datasets, and exhibits improved reproducibility.

bioinformatics↗

Systematic review and meta-analysis of protocols and yield of direct from sputum sequencing of Mycobacterium tuberculosis

Direct sputum whole genome sequencing (dsWGS) can revolutionize Mycobacterium tuberculosis (Mtb) diagnosis by enabling rapid detection of drug resistance and strain diversity without the biohazard of culture. We searched PubMed, Web of Science and Google scholar, and identified 8 studies that met inclusion criteria for testing protocols for dsWGS. Utilising meta-regression we identify several key factors positively associated with dsWGS success, including higher Mtb bacillary load, mechanical disruption, and enzymatic/chemical lysis. Specifically, smear grades of 3+ (OR = 14.7, 95% CI: 3.5, 62.1; p = 0.0005) were strongly associated with improved outcomes, whereas decontamination with sodium hydroxide (NaOH) was negatively associated (OR = 0.005, 95% CI: 0.001, 0.03; p = 7e-06), likely due to its harsh effects on Mtb cells. Furthermore, mechanical lysis (OR = 193.3, 95% CI: 11.7, 3197.8; p = 0.008) and enzymatic/chemical lysis (OR = 18.5, 95% CI: 1.9, 183.1; p = 0.02) were also strongly associated with improved dsWGS. Across the studies, we observed a high degree of variability in approaches to sputum pre-processing prior to dsWGS highlighting the need for standardized best practices. In particular we conclude that optimizing pre-processing steps including decontamination with the exploration of alternatives to NaOH to better preserve Mtb cells and DNA, and best practices for cell lysis during DNA extraction as priorities. Further and considering the strong association between Mtb load and successful dsWGS, protocol improvements for optimal sputum sample collection, handling, and storage could also further enhance the success rate of dsWGS.

microbiology↗