bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.01.06.697867

Home is where the host is: Evolutionary history of geographic spread, host switching, and adaptive genomic signatures in two generalist Group B Streptococcus clonal groups

Abstract

Group B Streptococcus (GBS) is a pathogen of global relevance in neonatal and maternal disease as well as bovine mastitis. Two closely related clonal groups, denoted 103 and 314 (CG103/314) have been detected in humans and cattle on multiple continents in recent decades but are poorly characterised compared to other host-generalist clades. We examined their potential origins, host-switching events and presence of a suite of genetic markers for antimicrobial resistance, virulence and host association using a newly assembled dataset of 248 CG103/314 genomes from humans, cattle, and food originating from five continents. We detected multiple host switches between humans and cattle, and significant regional differences in AMR gene distribution, possibly reflecting local differences in antimicrobial use across countries and hosts and indicating a capacity for regional adaptation to selective pressures. Across the evolutionary history of CG103/314 from both host species, the prevalence of the Lac.2 operon, a genetic marker associated with bovine host adaptation, was high, whereas the prevalence of the scpB-lmb gene pair, a genetic marker of human host adaptation in other GBS clonal groups, was very low. All isolates with scpB-lmb were associated with human disease rather than carriage. Our dataset displayed biases typical of research into multi-host pathogens, when sampling is often focused on a specific host species or setting. Consistent, balanced, contemporaneous and sympatric sampling efforts across host species and sources are needed for a full understanding of the distribution and emergence of CG103/314 and similar multi-host pathogens impacting food safety and public health. Impact statementThis study provides a comprehensive, global genomic overview of generalist clonal groups 103/314 of the human and animal pathogen Group B Streptococcus (GBS). By analysing host switching, antimicrobial resistance and virulence-associated markers, we show that these clonal groups display adaptation patterns shaped by region- and host-specific selective pressures. Our findings include potential expansion of the host range from humans and cattle into porcupines and pigs, and provides detailed discussion around anthropocentric sampling bias, highlighting the importance of balanced, multi-host sampling of generalist GBS lineages and One Health pathogens in general. This work reinforces the need for coordinated One Health surveillance to monitor emerging sub-lineages with relevance for food safety, human and animal health.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hilbig, A., Crestani, C., Barkham, T., Chen, S. L., Cobo-Angel, C., Ceballos-Marquez, A., Sirimanapong, W., Amin-Nordin, S., Nguyen, P. N., Castro Abreu Pinto, T., Andrade de Oliveira, L. M., Garbarino, C. A., Ricchi, M., Lembo, T., Lycett, S. J., Biek, R., Forde, T. L., Zadoks, R. N.. 2026-01-06. Home is where the host is: Evolutionary history of geographic spread, host switching, and adaptive genomic signatures in two generalist Group B Streptococcus clonal groups. https://doi.org/10.64898/2026.01.06.697867

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Ancestree: unified likelihood inference of ancestral alleles under supplied or inferred genealogies

Inferring ancestral states--determining, at each polymorphic site, which allele is ancestral and which derived--underpins many downstream population-genetic analyses, from selection scans and the unfolded site-frequency spectrum to demographic inference. However, no existing tool uniformly supports the full range of relevant inputs: plain variant data or ancestral recombination graphs (ARGs), with or without outgroups, while accommodating poly-allelic and recurrently-mutated sites. Here we present Ancestree, a likelihood-based engine that unifies these inputs within a single framework and returns full posteriors over the four nucleotide states at every site. It runs in three modes: a fixed-tree mode that assumes a single topology across sites and co-infers the per-branch substitution rates by maximum likelihood; an ARG mode that reads a different local tree at each site directly from a supplied ancestral recombination graph; and a local-tree mode that instead samples those local trees from the genotype data via a pairwise-coalescent HMM, needing no pre-existing ARG. On simulated data, the genealogy-based modes (ARG and local-tree) are more accurate and scale better, and remain robust under outgroup configurations that violate the fixed-tree assumption. Outgroups themselves remain difficult to replace: per-site inference accuracy on ingroup-polymorphic sites is markedly limited without them, and improves substantially with a single outgroup. The hardest sites are those fixed for the derived allele within the ingroup, which carry no within-ingroup signal and so need several sufficiently deep outgroups to recover, yet these are also highly informative downstream, carrying the high-frequency divergence signal on which selection and adaptation analyses often depend. Ancestree is available at github.com/Sendrowski/Ancestree.

evolutionary biology↗

Mapping tsetse fly connectivity in Uganda with machine learning landscape genetics

Introduction - Tsetse flies (genus Glossina) are biting insects that transmit human and animal trypanosomiases across sub-Saharan Africa, and sustainable vector control depends on understanding dispersal barriers and reinvasion routes. Despite major progress toward elimination, Uganda remains at risk for both human forms of the disease (Trypanosoma brucei gambiense and T. b. rhodesiense) and planners still lack reliable maps of tsetse movement and reinvasion risk. Methods and Results - We address this gap with machine-learning landscape genetics and species distribution models, integrating estimates of population genetic distance and geospatial environmental data to predict and map Glossina fuscipes fuscipes connectivity across Uganda and western Kenya. Inputs included microsatellite genotypes from 11 loci genotyped in 2,736 flies sampled from 87 localities and remotely sensed environmental predictors summarized along least-cost paths. Random forest models predicted patterns of genetic differentiation better than distance-only models, supporting the use of a machine-learning framework for connectivity inference across complex heterogeneous landscapes, and identified variables related to temperature and water availability as the strongest predictors of genetic connectivity. Conclusions - Combining landscape genetics predictions of connectivity with a species distribution model revealed regions with high habitat suitability but low connectivity that represent priority zones for area-wide integrated pest management strategies, including established riverine control tools such as tiny targets and other targeted interventions aimed at reducing reinvasion risk. These results provide biologically interpretable maps and quantitative uncertainty metrics that can guide targeted tsetse control, while providing a transferable analytical pipeline for modeling and mapping genetic connectivity across other species and landscapes.

evolutionary biology↗

Conserved genes with variable expression and function: Vasa, Piwi, and Dnmt1 in Oncopeltus fasciatus gametogenesis

The production of gametes is one of the universal processes of life and given millennia of evolution many of the core genes involved are expected to be highly refined and resistant to change. Here we examine the pattern of expression and functionally characterize three genes widely involved in gametogenesis: Vasa, Piwi, and Dnmt1. We examine their cellular location during both oogenesis and spermatogenesis in Oncopeltus fasciatus, a hemipteran insect, using fluorescent in situ hybridization chain reaction. In females, expression of Piwi was restricted to the trophocytes, but Vasa and Dnmt1 were expressed in developing ooctyes as well as the trophoctyes. In males, Piwi was expressed in the testis germline stem cells (GSCs), but Vasa and Dnmt1 expression was absent from these cells. Additionally, Piwi and Vasa were expressed in secondary spermatogonia and spermatocytes while Dnmt1 was predominantly expressed in the primary spermatogonia. We also functionally characterized their necessity for gametogenesis after RNAi-mediated gene expression knockdown. Vasa was not required for oogenesis but was required during spermatogenesis. Piwi and Dnmt1 were required for both oogenesis and spermatogenesis. These results were in contrast to their requirement for these processes from other organisms, which highlight the unexpected variation found in the processes and suggest that conservation depends on the level at which genes are examined: sequence, expression, or phenotypic and biochemical function. This suggests there is a more nuanced evolutionary story of conserved, yet plastic, gametogenic gene set in Metazoa that deserves further study across more organisms.

evolutionary biology↗