bioRxiv Science⌕ Search

Biology subjects

Barf, L.-M.

Publications and source records attributed to Barf, L.-M..

2 recordsLinked to original sources

Pangenome calculation beyond the species level using RIBAP: A comprehensive bacterial core genome annotation pipeline based on Roary and pairwise ILPs

Pangenome analysis is a computational method for identifying genes that are present or absent from a group of genomes, which helps to understand evolutionary relationships and to identify essential genes. While current state-of-the-art approaches for calculating pangenomes comprise various software tools and algorithms, these methods can have limitations such as low sensitivity, specificity, and poor performance on specific genome compositions. A common task is the identification of core genes, i.e., genes that are present in (almost) all input genomes. However, especially for species with high sequence diversity, e.g., higher taxonomic orders like genera or families, identifying core genes is challenging for current methods. We developed RIBAP (Roary ILP Bacterial core Annotation Pipeline) to specifically address these limitations. RIBAP utilizes an integer linear programming (ILP) approach that refines the gene clusters initially predicted by the pangenome pipeline Roary. Our approach performs pairwise all-versus-all sequence similarity searches on all annotated genes for the input genomes and translates the results into an ILP formulation. With the help of these ILPs, RIBAP has successfully handled the complexity and diversity of Chlamydia, Klebsiella, Brucella, and Enterococcus genomes, even when genomes of different species are part of the analysis. We compared the results of RIBAP with other established and recent pangenome tools (Roary, Panaroo, PPanGGOLiN) and showed that RIBAP identifies all-encompassing core gene sets, especially at the genus level. RIBAP is freely available as a Nextflow pipeline under the GPL3 license: https://github.com/hoelzer-lab/ribap.

bioinformatics↗

Extensive genomic divergence among 61 strains of Chlamydia psittaci

Chlamydia (C.) psittaci, the causative agent of avian chlamydiosis and human psittacosis, is a genetically heterogeneous species. Its broad host range includes parrots and many other birds, but occasionally also humans (via zoonotic transmission), ruminants, horses, swine and rodents. To assess whether there are genetic markers associated with host tropism we comparatively analyzed whole-genome sequences of 61 C. psittaci strains. Initially, poorly assembled genomes in public databases were subjected to clean-up, reassembly and polishing. Multiple sequence alignment of the genome sequences revealed four major clades within this species. Clade 1 represents the most recent lineage comprising 40/61 strains and contains 9/10 of the psittacine strains, including type strain 6BC, and 10/13 of human isolates. Clades 2-4 carry strains from different non-psittacine hosts. We found that clade membership correlates with classification schemes based on SNP types, ompA genotypes, multilocus sequence types as well as plasticity zone (PZ) structure and host preference. Genome analysis also revealed that i) sequence variation in the major outer membrane porin OmpA can result in 3D structural changes of immunogenic domains, ii) past host change of Clade 3 and 4 strains could be associated with loss of MAC/perforin in the PZ, rather than the large cytotoxin, iii) the distinct phylogeny of atypical strains (Clades 3 and 4) is also reflected in their repertoire of inclusion proteins (Inc family) and polymorphic membrane proteins (Pmps). Altogether, our study identified a number of genomic features that can be correlated with the phylogeny and host preference of C. psittaci strains.

genomics↗