bioRxiv Science⌕ Search

Biology subjects

Jeffrey, B. M.

Publications and source records attributed to Jeffrey, B. M..

2 recordsLinked to original sources

Gene conversion is a key driver of diversity hotspots in M. tuberculosis antigens and virulence-associated loci

Despite the long-held view of Mycobacterium tuberculosis (Mtb) as a genetically conserved pathogen, many genomic regions remain poorly resolved due to high sequence homology and repetitive content. Using complete genome assemblies generated from long-read sequencing of 151 globally representative clinical isolates, we comprehensively analyzed genome-wide patterns of genetic diversity and evolution across the Mtb genome. Our analysis uncovers pronounced diversity hotspots within paralogous regions generated by recurrent gene conversion between homologous genes. In many cases, these hotspots exhibit more than an order of magnitude greater genetic diversity than the rest of the Mtb genome, which is otherwise characterized by remarkably low variation. Mutations within these regions display clustered substitution patterns, excess paralog-matching variants, and distinct mutational spectra consistent with ongoing gene conversion. Our analysis identifies over 300 individual gene conversion events distributed throughout the Mtb phylogeny. These gene conversion events occur predominantly within gene families associated with virulence and host-pathogen interactions, including the PE, PPE, and ESX families. Several of the most pronounced diversity hotspots occur in antigens encoded within paralogous regions. Among these, the vaccine candidate PPE18 harbors mutations in validated epitope sequences and predicted alterations in HLA-II binding. Together, these findings demonstrate that gene conversion actively shapes antigenic and virulence-associated diversity in Mtb.

genomics↗

Analysis of the limited M. tuberculosis accessory genome reveals potential pitfalls of pan-genome analysis approaches

Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety of methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. To quantify sources of bias and error related to common pan-genome analysis approaches, we evaluated different approaches applied to curated collection of 151 Mycobacterium tuberculosis (Mtb) isolates. Mtb is characterized by its clonal evolution, absence of horizontal gene transfer, and limited accessory genome, making it an ideal test case for this study. Using a state-of-the-art graph-genome approach, we found that a majority of the structural variation observed in Mtb originates from rearrangement, deletion, and duplication of redundant nucleotide sequences. In contrast, we found that pan-genome analyses that focus on comparison of coding sequences (at the amino acid level) can yield surprisingly variable results, driven by differences in assembly quality and the softwares used. Upon closer inspection, we found that coding sequence annotation discrepancies were a major contributor to inflated Mtb accessory genome estimates. To address this, we developed panqc, a software that detects annotation discrepancies and collapses nucleotide redundancy in pan-genome estimates. When applied to Mtb and E. coli pan-genomes, panqc exposed distinct biases influenced by the genomic diversity of the population studied. Our findings underscore the need for careful methodological selection and quality control to accurately map the evolutionary dynamics of a bacterial species.

bioinformatics↗