bioRxiv Science⌕ Search

Biology subjects

Frankel, L. E.

Publications and source records attributed to Frankel, L. E..

2 recordsLinked to original sources

Low accuracy of complex admixture graph inference from f-statistics

F-statistics are commonly used to assess hybridization, admixture or introgression between populations or deeper evolutionary lineages. Using simulations, we find that network complexity had a large impact on the accuracy to infer the network structure from f statistics. Networks recovered accurately had one reticulation, or had their reticulations in "large" cycles of at least 4 nodes in all subnetworks. But accuracy was extremely poor to infer complex networks, in which a reticulation is part of a small cycle of only 3 nodes in some subnetwork. Accuracy also decreased with increasing number of reticulations and the network level. For these networks, accuracy was low even from large data sets with low mutation rate, under a molecular clock, and retaining many top-scoring graphs. Yet in all cases, the networks major tree was recovered reliably. When the molecular clock was violated, the f4-test tended to falsely detect the presence of reticulation in large data sets or under a high mutation rate. Rate variation also impacted network inference accuracy and increased the rate of falsely rejecting 1 reticulation as being adequate. We propose that identifiability, or lack thereof, is underlying the contrasting recoverability between simple and complex networks. Our findings suggest that the major tree is one feature that might be estimable from f-statistics. In practice, we recommend evaluating a large set of top-scoring networks inferred from f-statistics, and even so, using caution in assuming that the true network is part of this set. The extent of rate variation should be assessed in the system under study, especially at deeper time scales, or when using fast-evolving loci.

evolutionary biology↗

Summary tests of introgression are highly sensitive to rate variation across lineages

AO_SCPLOWBSTRACTC_SCPLOWThe evolutionary implications and frequency of hybridization and introgression are increasingly being recognized across the tree of life. To detect hybridization from multi-locus and genome-wide sequence data, a popular class of methods are based on summary statistics from subsets of 3 or 4 taxa. However, these methods often carry the assumption of a constant substitution rate across lineages and genes, which is commonly broken in many groups. In this work, we quantify the effects of rate variation on the D-statistic (also known as ABBA-BABA test), the D3 statistic, and HyDe. All three tests are used widely across a range of taxonomic groups, in part because they are very fast to compute. We consider rate variation across species lineages, across genes, their lineage-by-gene interaction, and rate variation across gene-tree edges. We simulated species networks according to a birth-death-hybridization process so as to capture a range of realistic species phylogenies. For all three methods tested, we found a marked increase in the false discovery of reticulation (type-1 error rate) when there is rate variation across species lineages. The D3 statistic was the most sensitive, with around 80% type-1 error, such that D3 appears to more sensitive to a departure from the clock than to the presence of reticulation. For all three tests, the power to detect hybridization events decreased as the number of hybridization events increased, indicating that multiple hybridization events can "hide" one another if they occur within a small subset of taxa. Our study highlights the need to consider rate variation when using site-based summary statistics, and points to the advantages of methods that do not require assumptions on evolutionary rates across lineages or across genes.

evolutionary biology↗