bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.07.13.738130

A hybrid approach combining a phylogenetic method and Approximate Bayesian Computation Random Forest for phylogenetic network inference: application to the rice domestication process in Asia

Abstract

Asian rice is one of the best documented crops in terms of genetic diversity. The domestication process, that probably started 9000 years ago in China, remains difficult to infer since the main vertical signal is blurred by horizontal signals related to gene flow among cultivars and wild relatives. Consequently, a large number of hypotheses on the domestication process of rice have been published. Besides, most of the methods used to infer these scenarios do not model all the known biological phenomena at stake. Here, we present a methodological study based on a rich stochastic model, that incorporates introgression events, incomplete lineage sorting, and mutations that happen over time. The global evolutionary scenario is represented by a phylogenetic network. Furthermore, each locus scenario is modeled according to a locus tree through the Multispecies Network Coalescent. More importantly, for inferring the phylogenetic network, we propose a new hybrid approach combining a phylogenetic network method and a machine learning technique. In particular, our hybrid approach, named SO_SCPLOWNARFC_SCPLOW, benefits from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW, and from the potential of a powerful machine learning classifier, i.e. Approximate Bayesian Computation Random Forest (ABC-RF). These two methods are complementary since SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW reconstructs network accurately, whereas ABC-RF is able to handle a large amount of data. The originality is twofold. First, prior distributions required for ABC-RF are calibrated thanks to SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOWs estimates. Secondly, ABC-RF relies on summary statistics inspired by phylogenetic network literature. We show, on simulated data, that the SO_SCPLOWNARFC_SCPLOW hybrid approach enjoys very good performances. On rice real data, it infers a scenario with a unique domestication (that of Japonica), followed by three reticulation events involving early Japonica. It highlights two introgression events at the origin of Indica and cAus, and one admixture event responsible for the emergence of cBas. Author summaryToday, in genomics, there is a real need for methods able to infer phylogenetic networks. A phylogenetic network is a directed graph representing events like hybridization, introgression, and horizontal gene transfer. Understanding these complex biological phenomena, essential for crop adaptation, can help breeders when facing challenges like climate change and population growth. Genome-wide diversity analysis thus requires network methods scaling for large data volumes and incorporating fundamental biological phenomena. In this context, we present a new hybrid approach, SO_SCPLOWNARFC_SCPLOW, that benefits from the potential of a powerful machine learning classifier, Approximate Bayesian Computation Random Forest, and from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW. Consequently, SO_SCPLOWNARFC_SCPLOW is able to handle large data-sets thanks to machine learning and is also based on a deep mathematical theory. On simulated data, our hybrid approach performs very well. When applied to real rice genomic data, it supports a scenario with a single domestication event, that of Japonica. The analysis further highlights the role of early Japonica in the origin of both Indica and circumAus. Finally, it identifies an ancient admixture event, involving circumAus in the emergence of circumBasmati. Together, these findings confirm the importance of early rice history along the Himalayan region.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rabier, C.-E., Berry, V., Glaszmann, J.-C.. 2026-07-17. A hybrid approach combining a phylogenetic method and Approximate Bayesian Computation Random Forest for phylogenetic network inference: application to the rice domestication process in Asia. https://doi.org/10.64898/2026.07.13.738130

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

RELAX does not reproduce its own estimates at default settings, and its output does not show it

Selection-intensity estimates from RELAX are reported as a point value of K with a likelihood-ratio P. We report that, at default settings and on data of ordinary size, the program does not reproduce its own fits. Of 27 enzyme entries refitted under two optimiser configurations, none reproduced its log-likelihood to within 0.01 units; the median change was 103 units, the largest over 3,400, and four verdicts reversed. Eighty null orthologues reproduced none. A byte-identical command returned a distinct likelihood on every repetition, single-threaded, across three releases, and on alignments simulated under the fitted model, where 3.3 per cent of replicates reproduced. The documented random-number seed never reaches the generator when assigned on the command line, yet reads back as the value supplied. PAML localises the cause: its two-ratio model, without site classes, reproduced its log-likelihood for all 288 genes; its site-class models agreed for 27 to 67 per cent. The instability follows the mixture over sites, not the program. The output does not show it: 46 of 410 fits ended with a negative likelihood-ratio statistic, impossible under convergence, and 123 of 410 report a K re-estimated under a domain restriction rather than the unconstrained maximum. Of 234 published studies using RELAX, none reported a seed. Seeding while holding the thread count at one reproduced sixty of sixty runs on twenty genes under two releases; the seed alone reproduced none of five, and no documentation states the second condition. We recommend that fits be repeated and their dispersion published.

evolutionary biology↗

Sequential accumulation of adaptive alleles forms an inversion supergene in deer mice

Supergenes are clusters of co-inherited loci that affect multiple or complex phenotypes. Despite the growing number of chromosomal inversions identified as supergenes in natural populations, their molecular basis and evolutionary history often remain obscure. Here, we identified two candidate genes, Slc45a2 and Npr3, within a 41-Mb inversion supergene in the deer mouse (Peromyscus maniculatus) that respectively drive darker coats and longer tails - two traits associated with forest adaptation. Mice homozygous for the inversion (inv/inv) exhibit elevated Slc45a2 expression in melanocytes relative to the congenic standard genotype (std/std), disrupting pheomelanin production. In parallel, downregulation of Npr3 in inv/inv mouse growth plates prolongs postnatal growth of caudal vertebrae, resulting in tail elongation. Population-level analyses further implicate that this supergene arose through the subsequent accumulation of the Npr3 allele within the inversion, rather than by capturing all beneficial mutations at its origin.

evolutionary biology↗

Toxin structure shapes palatability in a chemically defended butterfly

The toxicity of chemical defences is well studied, but the potential contribution of compound structure to predator deterrence remains largely unexplored. Whether predation acts more strongly on toxicity or unpalatability remains largely untested, partly because few systems allow toxin structure to vary independently of quantity. Heliconius sara larvae provide such a system: those reared on Passiflora auriculata sequester cyclopentenyl cyanogenic glucosides (CGs), while those reared on P. biflora biosynthesise comparable quantities of aliphatic CGs. Using two invertebrate predators, Camponotus floridanus ants and Hierodula membranacea mantids, we tested whether this structural difference affects palatability independent of toxicity. Mantids rejected larvae with cyclopentenyl CGs more often than larvae with aliphatic CGs, despite no detectable difference in total CG content. This pattern was mirrored in extract-based assays with ants, independently of cyanide release: extracts with cyclopentenyl CGs remained deterrent, while extracts with aliphatic CGs did not differ in deterrence from water. Live larvae, by contrast, elicited similar responses from ants regardless of CG structure. These results show that variation in toxin structure can strongly affect palatability, with some compounds conferring greater protection than others. This demonstrates the importance of chemical structural diversity in the evolution of chemical defences.

evolutionary biology↗