bioRxiv Science⌕ Search

Biology subjects

Harke, A. S.

Publications and source records attributed to Harke, A. S..

2 recordsLinked to original sources

PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining

The exponential growth of microbial genome data presents unprecedented opportunities for mining the potential of microorganisms. The burgeoning field of pangenomics offers a framework for extracting insights from this big biological data. Recent advances in microbial pangenomic research have generated substantial data and literature, yielding valuable knowledge across diverse microbial species. PanKB (pankb.org), a knowledgebase designed for microbial pangenomics research and biotechnological applications, was built to capitalize on this wealth of information. PanKB currently includes 51 pangenomes on 8 industrially relevant microbial families, comprising 8, 402 genomes, over 500, 000 genes, and over 7M mutations. To describe this data, PanKB implements four main components: 1) Interactive pangenomic analytics to facilitate exploration, intuition, and potential discoveries; 2) Alleleomic analytics, a pangenomic- scale analysis of variants, providing insights into intra-species sequence variation and potential mutations for applications; 3) A global search function enabling broad and deep investigations across pangenomes to power research and bioengineering workflows; 4) A bibliome of 833 open- access pangenomic papers and an interface with an LLM that can answer in-depth questions using their knowledge. PanKB empowers researchers and bioengineers to harness the full potential of microbial pangenomics and serves as a valuable resource bridging the gap between pangenomic data and practical applications. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=134 SRC="FIGDIR/small/608241v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@4d3daaorg.highwire.dtl.DTLVardef@10b6d49org.highwire.dtl.DTLVardef@133da4forg.highwire.dtl.DTLVardef@141ae7e_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

Genomic insights into Lactobacillaceae: Analyzing the Alleleome of core pangenomes for enhanced understanding of strain diversity and revealing Phylogroup-specific unique variants

The Lactobacillaceae familys significance in food and health, combined with available strain-specific genomes, enables genome assessment through pangenome analysis. The Alleleome of the core pangenomes of the Lactobacillaceae family, which identifies natural sequence variations, was reconstructed from the amino acid and nucleotide sequences of the core genes across 2,447 strains of 26 species. It comprised 3.71 million amino acid variants in 29,448 core genes across the family. The alleleome analysis of the Lactobacillaceae family revealed key findings: 1) In the core pangenome, amino acid substitutions prevailed over rare insertions and deletions, 2) Purifying negative selection primarily influenced core gene variations in the family, with diversifying selection noted in L. helveticus. L. plantarums core alleleome was investigated due to its industrial importance. In L. plantarum, the defining characteristics of its core alleleome included: 1) It is highly conserved; 2) Among 235 isolation sources, the primary categories displaying variant prevalence were fermented food, feces, and unidentified sources; 3) It is predominantly characterized by conservative and moderately conservative mutations; and 4) Phylogroup-specific core variant gene analysis identified unique variants (DltX, FabZ1, Pts23B, CspP) in phylogroups I and B which could be used as identifier or validation markers of strain or phylogroup.

genomics↗