bioRxiv Science⌕ Search

Biology subjects

Altinkaya, I.

Publications and source records attributed to Altinkaya, I..

3 recordsLinked to original sources

vcfgl: A flexible genotype likelihood simulator for VCF/BCF files

MotivationAccurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. ResultsWe present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, VCF/BCF and gVCF file formats. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. Availabilityvcfgl is freely available at https://github.com/isinaltinkaya/vcfgl. Contactisin.altinkaya@sund.ku.dk Supplementary informationSupplementary information is available online.

bioinformatics↗

Steppe Ancestry in western Eurasia and the spread of the Germanic Languages

Today, Germanic languages, including German, English, Frisian, Dutch and the Nordic languages, are widely spoken in northwest Europe. However, key aspects of the assumed arrival and diversification of this linguistic group remain contentious1-3. By adding 712 new ancient human genomes we find an archaeologically elusive population entering Sweden from the Baltic region by around 4000 BP. This population became widespread throughout Scandinavia by 3500 BP, matching the contemporaneous distribution of Palaeo-Germanic, the Bronze Age predecessor of Proto-Germanic4-6. These Baltic immigrants thus offer a new potential vector for the first Germanic speakers to arrive in Scandinavia, some 800 years later than traditionally assumed7-12. Following the disintegration of Proto-Germanic13-16, we find by 1650 BP a southward push from Southern Scandinavia into presumed Celtic-speaking areas, including Germany, Poland and the Netherlands. During the Migration Period (1575-1375 BP), we see this ancestry representing West Germanic Anglo-Saxons in Britain, and Langobards in southern Europe. We find a related large-scale northward migration into Denmark and South Sweden corresponding with historically attested Danes and the expansion of Old Norse. These movements have direct implications for multiple linguistic hypotheses. Our findings show the power of combining genomics with historical linguistics and archaeology in creating a unified, integrated model for the emergence, spread and diversification of a linguistic group.

genetics↗

HMMploidy: inference of ploidy levels from short-read sequencing data

The inference of ploidy levels from genomic data is important to understand molecular mechanisms underpinning genome evolution. However, current methods based on allele frequency and sequencing depth variation do not have power to infer ploidy levels at low-and mid-depth sequencing data, as they do not account for data uncertainty. Here we introduce HMMploidy, a novel tool that leverages the information from multiple samples and combines the information from sequencing depth and genotype likelihoods. We demonstrate that HMMploidy outperforms existing methods in most tested scenarios, especially at low-depth with large sample size. We apply HMMploidy to sequencing data from the pathogenic fungus Cryptococcus neoformans and retrieve pervasive patterns of aneuploidy, even when artificially downsampling the sequencing data. We envisage that HMMploidy will have wide applicability to low-depth sequencing data from polyploid and aneuploid species.

bioinformatics↗