bioRxiv Science⌕ Search

Biology subjects

Heinzel, C. S.

Publications and source records attributed to Heinzel, C. S..

3 recordsLinked to original sources

Enhancing Intra-Continental Biogeographical Ancestry Prediction Through a Machine Learning Marker Selection Method

While classifiers such as TabPFN (Hollmann et al., 2025) and SNIPPER (Phillips et al., 2007a) achieve strong intercontinental performance (Heinzel et al., 2025), their accuracy in classifying individuals within Europe remains low. One major factor contributing to this limitation is the set of genetic markers used for classification. Marker panels such as the VISAGE Enhanced Tool (Xavier et al., 2022) are commonly employed in forensic genetics because they contain ancestry-informative markers (AIMs) that distinguish very well between major continental populations. However, these panels are often not optimized for fine-scale differentiation within continents, where genetic variation is more subtle and population structure is rather continuous. We apply machine learning to select informative markers for intra-European classification, using data from Consortium et al. (2015). Compared with the VISAGE Enhanced Tool and allele frequency-based approaches (Phillips et al., 2007b; Kosoy et al., 2009; Nassir et al., 2009; Kidd et al., 2014; Phillips et al., 2014a), our marker sets achieve substantially higher accuracy within Europe: For four European populations, accuracy improves from 68.2% (VISAGE, 104 markers) to 73.7% (100 new markers) and 82.3% (200 new markers). For five populations, accuracy rises from 56.1% (VISAGE) to 64.5% (100 new markers). Our results show that tailored marker selection markedly improves intra-continental classification. While optimized here for Europe, the method can be applied to any region with sufficient training data.

genetics↗

Advancing Biogeographical Ancestry Predictions Through Machine Learning

Tools like Snipper or the Admixture Model count as state-of-the-art methods in forensic science for biogeographical ancestry. However, they have not been systematically compared to classifiers widely used in other disciplines. Noting that genetic data have a tabular form, this study addresses this gap by benchmarking forensic classifiers against TabPFN, a cutting-edge, general-purpose machine learning classifier for tabular data. The comparison evaluates performance using metrics such as accuracy--the proportion of correct classifications--and ROC AUC. We examine classification tasks for individuals at both the intracontinental and continental levels, based on a published dataset for training and testing. Our results reveal significant performance differences between methods, with TabPFN consistently achieving the best results for accuracy, ROC AUC and log loss. E.g., for accuracy, TabPFN improves SNIPPER from 84% to 93% on a continental scale using eight populations, and from 43% to 48% for inter-European classification with ten populations.

genetics↗

Revealing the range of maximum likelihood estimates in the admixture model.

Many ancestry inference tools, including STRUCTURE and ADMIXTURE, rely on the admixture model to infer both, allele frequencies p and individual admixture proportions q for a collection of individuals relative to a set of hypothetical ancestral populations. We show that under realistic conditions the likelihood in the admixture model is typically flat in some direction around a maximum likelihood estimate [Formula]. In particular, the maximum likelihood estimator is non-unique and there is a complete spectrum of possible estimates. Common inference tools typically identify only a few points within this spectrum. We provide an algorithm which computes the set of equally likely [Formula], when starting from [Formula]. It is analytic for K = 2 ancestral populations and numeric for K > 2. We apply our algorithm to data from the 1000 genomes project, and show that inter-European estimators of q can come with a large set of equally likely possibilities. In general, markers with large allele frequency differences between populations in combination with individuals with concentrated admixture proportions lead to small areas with a flat likelihood. Our findings imply that care must be taken when interpreting results from STRUCTURE and ADMIXTURE if populations are not separated well enough.

genetics↗