Proteogenomic discovery of novel small proteins in clinical Mycobacterium tuberculosis strains
Even though our meta-analysis ranks Mycobacterium tuberculosis genomes among the bacterial pathogens that are most straightforward to assemble, most available assemblies relied on short-read sequencing and contain genomic blind spots that miss functionally important genes. Complete genomes are essential for functional genomics, particularly for identifying small ORF-encoded proteins (SEPs; [≤]100 amino acids), which can play critical biological roles yet are frequently missed by standard annotations. Here, we generated complete long-read assemblies for six clinical reference strains representing lineage 1 and the more pathogenic lineage 2, followed by comparative genomic and proteogenomic analyses. We additionally provide software to predict comprehensive sets of mycobacteria-specific proline-glutamic acid (PE) and PPE family genes, including lineage-specific variants. Using parallel accumulation-serial fragmentation mass spectrometry, we detected approximately two-thirds of each strains annotated proteome from unfractionated cell extracts. Extending our proteogenomic framework across related strains, and adding rigorous control of proteogenomic discovery rates using entrapment strategies, we revealed 12-24 previously unannotated proteins per strain, predominantly SEPs, 56-60 alternative translation start sites, and 9-17 expressed pseudogenes. Newly identified proteins included conserved and lineage-specific SEPs, an antitoxin, candidate antimicrobial peptides and novel proteins under purifying selection. Overall, applying this improved proteogenomics method to phylogenomically selected clinical reference strains provides a valuable approach for discovering candidate diagnostics or therapeutics, as illustrated here for a WHO-listed critical bacterial pathogen.