bioRxiv Science⌕ Search

Biology subjects

Tallman, S.

Publications and source records attributed to Tallman, S..

4 recordsLinked to original sources

Genomic variation and ancestry of Helicobacter pylori in the admixed population of Cabo Verde reveal host adaptation, limited host-pathogen ancestry concordance, and signatures of trans-Atlantic slave trade migrations

The genomic structure of the stomach-colonising, gastric cancer-associated bacterium Helicobacter pylori mirrors human evolutionary history, making it difficult to disentangle shared ancestry from true host-pathogen interactions driving disease risk. Recently admixed populations offer opportunities for new coevolutionary trajectories between H. pylori and its host that may disrupt parallel evolution. Here, we analyse host H. pylori IgG, blood serum biomarkers, and paired host-bacterial genomic data from the general population (independent of dyspepsia symptoms) of Cabo Verde. Seropositivity is high (82.8%), as in neighbouring African populations. Serum biomarkers indicate that this unique cohort mainly comprises hosts with limited gastric inflammation. Genomic analyses show that Cabo Verdean H. pylori derive European (hspSWeurope) and West African (hspAfrica1WAfrica) ancestry. The delineation of a reference dataset that includes new strains from Ghana and Portugal, uncovered a distinct West-Central African cluster (hspAfrica1WCAfrica). This dataset allowed inference of within-Africa and within-Europe ancestries in Cabo Verdean and American strains that reflect historical migrations linked to European colonisation and the trans-Atlantic Slave Trade. European strains in Cabo Verde show additional population structure, including two low-diversity clusters. Approximate Bayesian Computation (ABC) inference indicates that the most common cluster represents a clonal expansion, likely driven by host adaptation leading to increased transmission within Cabo Verde. Notably, these strains lack known virulence genes and show significant divergence from gastric cancer strains. Finally, host and bacterial ancestry are uncoupled, indicating disruption of ancestral parallel evolution. Cabo Verde, therefore, provides a unique natural model for studying host-pathogen interactions underlying susceptibility to H. pylori.

evolutionary biology↗

Tracing the evolutionary histories of ultra-rare variants using variational dating of large ancestral recombination graphs

Ultra-rare variants dominate whole-genome sequencing datasets, yet their interpretation is limited by allele frequency, which provides little information at very low counts and is highly sensitive to uneven ancestry representation. Allele age offers an ancestry-agnostic alternative but existing methods do not scale to biobank-sized cohorts. Here we present a scalable variational algorithm for dating Ancestral Recombination Graphs (ARGs), implemented in tsdate, together with new distributed methods enabling practical biobank-scale ARG inference using tsinfer. Applied to 47,535 genomes from the Genomics England 100,000 Genomes Project, we infer contiguous ARGs spanning 206 Mb and estimate ages for 23.2 million variants, including 11.8 million singletons. ARG-based allele ages remain accurate under extreme sampling imbalance and, in real data, reveal signatures of purifying selection and clinically relevant heterogeneity among variants with identical observed frequencies. Estimates for recent mutations are precise only at large sample sizes, highlighting the information accessible in the haplotype structure of large datasets. Biobank-scale ARGs therefore enable robust, ancestry-agnostic age estimation for ultra-rare variation with broad utility for statistical and clinical genomics.

genetics↗

Analysis-ready VCF at Biobank scale using Zarr

BackgroundVariant Call Format (VCF) is the standard file format for interchanging genetic variation data and associated quality control metrics. The usual row-wise encoding of the VCF data model (either as text or packed binary) emphasises efficient retrieval of all data for a given variant, but accessing data on a field or sample basis is inefficient. Biobank scale datasets currently available consist of hundreds of thousands of whole genomes and hundreds of terabytes of compressed VCF. Row-wise data storage is fundamentally unsuitable and a more scalable approach is needed. ResultsZarr is a format for storing multi-dimensional data that is widely used across the sciences, and is ideally suited to massively parallel processing. We present the VCF Zarr specification, an encoding of the VCF data model using Zarr, along with fundamental software infrastructure for efficient and reliable conversion at scale. We show how this format is far more efficient than standard VCF based approaches, and competitive with specialised methods for storing genotype data in terms of compression ratios and single-threaded calculation performance. We present case studies on subsets of three large human datasets (Genomics England: n=78,195; Our Future Health: n=651,050; All of Us: n=245,394) along with whole genome datasets for Norway Spruce (n=1,063) and SARS-CoV-2 (n=4,484,157). We demonstrate the potential for VCF Zarr to enable a new generation of high-performance and cost-effective applications via illustrative examples using cloud computing and GPUs. ConclusionsLarge row-encoded VCF files are a major bottleneck for current research, and storing and processing these files incurs a substantial cost. The VCF Zarr specification, building on widely-used, open-source technologies has the potential to greatly reduce these costs, and may enable a diverse ecosystem of next-generation tools for analysing genetic variation data directly from cloud-based object stores, while maintaining compatibility with existing file-oriented workflows. Key PointsO_LIVCF is widely supported, and the underlying data model entrenched in bioinformatics pipelines. C_LIO_LIThe standard row-wise encoding as text (or binary) is inherently inefficient for large-scale data processing. C_LIO_LIThe Zarr format provides an efficient solution, by encoding fields in the VCF separately in chunk-compressed binary format. C_LI

bioinformatics↗

Long-term evolution of Streptococcus mitis and Streptococcus pneumoniae leads to higher genetic diversity within rather than between human populations

Evaluation of the apportionment of genetic diversity of bacterial commensals within and between populations is an important step in the characterization of their evolutionary potential. Recent studies showed a correlation between the genomic diversity of human commensal strains and that of their host, but the strength of this correlation and of the geographic structure among populations is a matter of debate. Here, we studied the genomic diversity and evolution of the phylogenetically related oronasopharingeal healthy-carriage Streptococcus mitis and Streptococcus pneumoniae, whose lifestyles range from stricter commensalism to high pathogenic potential. A total of 119 S. mitis genomes showed higher within- and among-host variation than 810 S. pneumoniae genomes in European, East Asian and African populations. Summary statistics of the site-frequency spectrum for synonymous and non-synonymous variation and ABC modelling showed this difference to be due to higher historical population effective size (Ne) in S. mitis, whose genomic variation been maintained close to mutation-drift equilibrium across (at least many) generations, whereas S. pneumoniae has been expanding from a smaller ancestral population and has been subjected to adaptive selection. Strikingly, both species show limited differentiation among populations. As genetic differentiation is inversely proportional to the product of effective population size and migration rate (Nem), we argue that large Ne have led to similar differentiation patterns, even if m is very low for S. mitis. We conclude that more diversity within than among human populations and limited population differentiation must be common features of the human microbiome due to large Ne.

evolutionary biology↗