bioRxiv Science⌕ Search

Biology subjects

Popli, D.

Publications and source records attributed to Popli, D..

3 recordsLinked to original sources

A joint framework for studying population structure using principal component analysis and F-statistics

Principal component analysis (PCA) and F-statistics are routinely used in population genetic and archaeogenetic studies. Here, we present a statistical framework to combine them into a joint analysis, showing where they coincide, and where slightly different assumptions made can lead to different outcomes. In particular, we discuss the differences of probabilistic PCA, Latent Subspace Estimation and classical PCA, and show that F-statistics are more naturally interpreted in a probabilistic PCA framework. We also show that individual-based F-statistics can be accurately estimated from probabilistic PCA in the presence of large amounts of missing data. We compare estimates from probabilistic PCA-based framework to ADMIXTOOLS 2 using simulations and published data, and show that this joint estimation framework addresses limitations of estimating F-statistics and PCA independently.

evolutionary biology↗

BREADR: An R Package for the Bayesian Estimation of Genetic Relatedness from Low-coverage Genotype Data

Robust and reliable estimates of how individuals are biologically related to each other are a key source of information when reconstructing pedigrees. In combination with contextual data, reconstructed pedigrees can be used to infer possible kinship practices in prehistoric populations. However, standard methods to estimate biological relatedness from genome sequence data cannot be applied to low coverage sequence data, such as are common in ancient DNA (aDNA) studies. Critically, a statistically robust method for assessing and quantifying the confidence of a classification of a specific degree of relatedness for a pair of individuals, using low coverage genome data, is lacking. In this paper we present the R-package BREADR (Biological RElatedness from Ancient DNA in R), which leverages the so-called pairwise mismatch rate, calculated on optimally-thinned genome-wide pseudo-haploid sequence data, to estimate genetic relatedness up to the second degree, assuming an underlying binomial distribution. BREADR also returns a posterior probability for each degree of relatedness, from identical twins/same individual, first-degree, second-degree or "unrelated" pairs, allowing researchers to quantify and report the uncertainty, even for particularly low-coverage data. We show that this method accurately recovers degrees of relatedness for sequence data with coverage as low as 0.04X using simulated data, and then compare the performance of BREADR on empirical data from Bronze Age Iberian human sequence data. The BREADR package is designed for pseudo-haploid genotype data, common in aDNA studies.

bioinformatics↗

KIN: A method to infer relatedness from low-coverage ancient DNA

Genetic kinship of ancient individuals can provide insights into their culture and social hierarchy, and is relevant for downstream genetic analyses. However, estimating relatedness from ancient DNA is difficult due to low-coverage, ascertainment bias, or contamination from various sources. Here, we present KIN, a method to estimate the relatedness of a pair of individuals from the identical-by-descent segments they share. KIN accurately classifies up to 3rd-degree relatives using [≥] 0.05x sequence coverage and differentiates siblings from parent-child. It incorporates additional models to adjust for contamination and detect inbreeding, which improves classification accuracy.

genetics↗