bioRxiv Science⌕ Search

Biology subjects

Güler, M. N.

Publications and source records attributed to Güler, M. N..

5 recordsLinked to original sources

Bias in genome-wide association test statistics due to omitted interactions

Over the past two decades, genome-wide association studies (GWAS) enabled the discovery of thousands of variants associated with many complex human traits. However, conventional GWAS are still widely performed with linear models with the assumption that the genetic effects are predominantly additive. In this work, we investigate the test statistic behavior when linear models are used to obtain significant genotype-phenotype associations without accounting for epistasis. We first algebraically derive mean and variance shift in the null statistic due to the omitted interaction term, and define the boundary between conservative (i.e., deflated statistic tail) and anti-conservative (i.e., inflated statistic tail) regimes for the common GWAS significance threshold. We then perform phenotype simulation analyses using the Estonian Biobank genotypes and validate the mathematical model. We demonstrate that the anti-conservative regime is plausible under realistic parameter settings and models omitting interaction terms can produce spurious significance. Our findings suggest caution when interpreting statistically significant signals reported in the literature based on linear models, especially for large-scale GWAS.

bioinformatics↗

An age-specific burial practice reflected in ancient DNA preservation in Neolithic Catalhöyük

Selective funerary practices can inform about social relationships in prehistoric societies but are often difficult to discern. Here we present evidence for an age-specific practice at the Neolithic site of Catalhoyuk in Anatolia, dating to the 7th millennium BCE. Among ancient DNA libraries produced from 362 petrous bone samples, those of subadults contained three times higher average human DNA than those of adults. This difference in organic preservation was also confirmed by FTIR analysis. Studying similar datasets from seven prehistoric and historical sites, we found a similar age-related difference in only one cemetery. We propose that the organic preservation difference with age was caused by the special treatment of chosen corpses before interment, such as defleshing or drying, which was more frequently applied to Catalhoyuk adults and promoted organic decay.

evolutionary biology↗

READv2: Advanced and user-friendly detection of biological relatedness in archaeogenomics

The possibility to obtain genome-wide ancient DNA data from multiple individuals has facilitated an unprecedented perspective into prehistoric societies. Studying biological relatedness in these groups requires tailored approaches for analyzing ancient DNA due to its low coverage, post-mortem damage, and potential ascertainment bias. Here we present READv2 (Relatedness Estimation from Ancient DNA version 2), an improved Python 3 re-implementation of the most widely used tool for this purpose. While providing increased portability and making the software future-proof, we are also able to show that READv2 (a) is orders of magnitude faster than its predecessor; (b) has increased power to detect pairs of relatives using optimized default parameters; and, when the number of overlapping SNPs is sufficient, (c) can differentiate between full-siblings and parent-offspring, and (d) can classify pairs of third-degree relatedness. We further use READv2 to analyze a large empirical dataset that has previously needed two separate tools to reconstruct complex pedigrees. We show that READv2 yields results and precision similar to the combined approach but is faster and simpler to run. READv2 will become a valuable part of the archaeogenomic toolkit in providing an efficient and user-friendly classification of biological relatedness from pseudohaploid ancient DNA data.

bioinformatics↗

Population genomic history of the endangered Anatolian and Cyprian mouflons in relation to worldwide wild, feral and domestic sheep lineages

Once widespread in their homelands, Anatolian mouflon (Ovis gmelini anatolica) and Cyprian mouflon (Ovis gmelini ophion) were driven to near extinction during the 20th century and are currently listed as endangered populations by the IUCN. While the exact origins of these lineages remain unclear, they have been suggested to be close relatives of domestic sheep or remnants of proto-domestic sheep groups. Here, we study whole genome sequences of n=5 Anatolian mouflons and n=10 Cyprian mouflons in terms of population history and diversity, relative to eight other extant sheep lineages. We find reciprocal genetic affinity between Anatolian and Cyprian mouflons and domestic sheep, higher than all other studied wild sheep genomes, including the Iranian mouflon (Ovis gmelini). Despite similar recent population dynamics, Anatolian and Cyprian mouflons exhibit disparate diversity levels, which can potentially be attributed to founder effects, island isolation, introgression from domestic lineages, or different bottleneck dynamics. The lower relative mutation load found in Cyprian compared to Anatolian mouflons suggests the purging of recessive deleterious variants in the former. This agrees with estimates of a long-term small effective population size in the Cyprian mouflon. Both subspecies harbor considerable numbers of runs of homozygosity (ROH) blocks <2 Mb, which reflects the effect of small population size. Expanding our analyses to worldwide wild and feral Ovis genomes, we observe varying viability metrics among different lineages, and a limited consistency between viability metrics and conservation status. Factors such as recent inbreeding, introgression, and unique population dynamics may contribute to the observed disparities.

evolutionary biology↗

Benchmarking kinship estimation tools for ancient genomes using pedigree simulations

There is growing interest in uncovering genetic kinship patterns in past societies using low-coverage paleogenomes. Here, we benchmark four tools for kinship estimation with such data: lcMLkin, NgsRelate, KIN, and READ, which differ in their input, IBD-estimation methods and statistical approaches. We used pedigree and ancient genome sequence simulations to evaluate these tools when only a limited number (1K to 50K) of shared SNPs (with minor allele frequency [&ge;]0.01) are available. The performance of all four tools was comparable using [&ge;]20K SNPs. We found that first-degree related pairs can be accurately classified even with 1K SNPs, with 85% F1 scores using READ and 96% using NgsRelate or lcMLkin. Distinguishing third-degree relatives from unrelated pairs or second-degree relatives was also possible with high accuracy (F1 >90%) with 5K SNPs using NgsRelate and lcMLkin, while READ and KIN showed lower success (69% and 79%, respectively). Meanwhile, noise in population allele frequencies and inbreeding (first cousin mating) led to deviations in kinship coefficients, with different sensitivities across tools. We conclude that using multiple tools in parallel might be an effective approach to achieve robust estimates on ultra-low coverage genomes.

genetics↗