bioRxiv Science⌕ Search

Biology subjects

Williams, M. P.

Publications and source records attributed to Williams, M. P..

5 recordsLinked to original sources

Optimised in-solution enrichment of over a million ancient human SNPs

In-solution hybridisation enrichment of genetic markers is a method of choice in paleogenomic studies, where the DNA of interest is generally heavily fragmented and contaminated with environmental DNA, and where the retrieval of genetic data comparable between individuals is challenging. Here, we benchmarked the commercial "Twist Ancient DNA" reagent from Twist Biosciences using sequencing libraries from ancestrally diverse ancient human samples with low to high endogenous DNA content (0.1-44%). For each library, we tested one and two rounds of enrichment, and assessed performance compared to deep shotgun sequencing. We find that the "Twist Ancient DNA" assay provides robust enrichment of [~]1.2M target SNPs without introducing allelic bias that may interfere with downstream population genetics analyses. Additionally, we show that pooling up to 4 sequencing libraries and performing two rounds of enrichment is both reliable and cost-effective for libraries with less than 27% endogenous DNA content. Above 38% endogenous content, a maximum of one round of enrichment is recommended for cost-effectiveness and to preserve library complexity. In conclusion, we provide researchers in the field of human paleogenomics with a comprehensive understanding of the strengths and limitations of different sequencing and enrichment strategies, and our results offer practical guidance for optimising experimental protocols.

genomics↗

The European Neolithic Expansion: A Model Revealing Intense Assortative Mating and Restricted Cultural Transmission

The Neolithic revolution initiated a pivotal change in human society, marking the shift from foraging to farming. The underlying mechanisms of agricultural expansion are debated, primarily between cultural diffusion (knowledge and practices transfer) and demic diffusion, or people migration and replacement. Ancient DNA analyses reveal significant ancestry changes during Europes Neolithic transition, suggesting primarily demic expansion. However, the presence of 10-15% hunter-gatherer ancestry in modern Europeans indicates cultural transmission and non-assortative mating were additional contributing factors. We integrate mathematical models, agent-based simulations, and ancient DNA analysis to dissect and quantify the roles of cultural diffusion and assortative mating in farmings expansion. Our findings indicate limited cultural transmission and predominantly within-group mating. Additionally, we challenge the assumption that demic spread always leads to ancestry turnover. This underscores the need to reassess prehistoric cultural expansions and offers new insights into early agricultural society through the integration of ancient DNA with archaeological models.

genomics↗

Testing Times: Challenges in Disentangling Admixture Histories in Recent and Complex Demographies

Paleogenomics has expanded our knowledge of human evolutionary history. Since the 2020s, the study of ancient DNA has increased its focus on reconstructing the recent past. However, the accuracy of paleogenomic methods in answering questions of historical and archaeological importance amidst the increased demographic complexity and decreased genetic differentiation within the historical period remains an open question. We used two simulation approaches to evaluate the limitations and behavior of commonly used methods, qpAdm and the f3-statistic, on admixture inference. The first is based on branch-length data simulated from four simple demographic models of varying complexities and configurations. The second, an analysis of Eurasian history composed of 59 populations using whole-genome data modified with ancient DNA conditions such as SNP ascertainment, data missingness, and pseudo-haploidization. We show that under conditions resembling historical populations, qpAdm can identify a small candidate set of true sources and populations closely related to them. However, in typical ancient DNA conditions, qpAdm is unable to further distinguish between them, limiting its utility for resolving fine-scaled hypotheses. Notably, we find that complex gene-flow histories generally lead to improvements in the performance of qpAdm and observe no bias in the estimation of admixture weights. We offer a heuristic for admixture inference that incorporates admixture weight estimate and P-values of qpAdm models, and f3-statistics to enhance the power to distinguish between multiple plausible candidates. Finally, we highlight the future potential of qpAdm through whole-genome branch-length f2-statistics, demonstrating the improved demographic inference that could be achieved with advancements in f-statistic estimations.

genomics↗

Allelic bias when performing in-solution enrichment of ancient human DNA

In-solution hybridisation enrichment of genetic variation is a valuable methodology in human paleogenomics. It allows enrichment of endogenous DNA by targeting genetic markers that are comparable between sequencing libraries. Many studies have used the 1240k reagent--which enriches 1,237,207 genome-wide SNPs--since 2015, though access was restricted. In 2021, Twist Biosciences and Daicel Arbor Biosciences independently released commercial kits that enabled all researchers to perform enrichments for the same 1240k SNPs. We used the Daicel Arbor Biosciences Prime Plus kit to enrich 132 ancient samples from three continents. We identified a systematic assay bias that increases genetic similarity between enriched samples and that cannot be explained by batch effects. We present the impact of the bias on population genetics inferences (e.g., Principal Components Analysis, [f]-statistics) and genetic relatedness (READ). We compare the Prime Plus bias to that previously reported of the legacy 1240k enrichment assay. In [f]-statistics, we find that all Prime-Plus-generated data exhibit artefactual excess shared drift, such that within-continent relationships cannot be correctly determined. The bias is more subtle in READ, though interpretation of the results can still be misleading in specific contexts. We expect the bias may affect analyses we have not yet tested. Our observations support previously reported concerns for the integration of different data types in paleogenomics. We also caution that technological solutions to generate 1240k data necessitate a thorough validation process before their adoption in the paleogenomic community.

genomics↗

False discovery rates of qpAdm-based screens for genetic admixture

qpAdm is a statistical tool that is often used for testing large sets of alternative admixture models for a target population. Despite its popularity, qpAdm remains untested on two-dimensional stepping-stone landscapes and in situations with low pre-study odds (low ratio of true to false models). We tested high-throughput qpAdm protocols with typical properties such as number of source combinations per target, model complexity, model feasibility criteria, etc. Those protocols were applied to admixture-graph-shaped and stepping-stone simulated histories sampled randomly or systematically. We demonstrate that false discovery rates of high-throughput qpAdm protocols exceed 50% for many parameter combinations since: 1) pre-study odds are low and fall rapidly with increasing model complexity; 2) complex migration networks violate the assumptions of the method, hence there is poor correlation between qpAdm p-values and model optimality, contributing to low but non-zero false positive rate and low power; 3) although admixture fraction estimates between 0 and 1 are largely restricted to symmetric configurations of sources around a target, a small fraction of asymmetric highly non-optimal models have estimates in the same interval, contributing to the false positive rate. We also re-interpret large sets of qpAdm models from two studies in terms of source-target distance and symmetry and suggest improvements to qpAdm protocols: 1) temporal stratification of targets and proxy sources in the case of admixture-graph-shaped histories; 2) focused exploration of few models for increasing pre-study odds; dense landscape sampling for increasing power and stringent conditions on estimated admixture fractions for decreasing the false positive rate. Article SummaryProliferation in the archaeogenetic literature of protocols for detection of admixed groups based a so-called qpAdm algorithm became disconnected from performance testing: the only extensive study of qpAdm on simulated data showed that it performs well under an unrealistically simple demographic scenario. We found that false discoveries of gene flows by qpAdm on a collection of random admixture-graph-shaped histories and on complex stepping-stone landscapes are very common and provide guidelines for design of qpAdm protocols in archaeogenetic studies.

genetics↗