bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.04.03.646226

A high-resolution genomic study of the Pama-Nyungan speaking Yolngu people of northeast Arnhem Land, Australia.

Abstract

ObjectivesAbout 300 Aboriginal languages were spoken in Australia. These were classified into two groups: Pama-Nyungan (PN), comprised of one language Family, and Non-Pama-Nyungan (NPN) with more than 20 language Families. The Yolngu people belong to the larger PN Family and live in Arnhem Land in northern Australia. They are surrounded by groups who speak NPN languages. This study, using nuclear genomic and mitochondrial DNA data, was undertaken to shed light on the origins of the Yolngu people and their language. The nuclear genomic sequences of Yolngu people were compared to those of other Indigenous Australians, as well as Papuan, African, East Asian and European people. Materials and methodsWith the agreement of Indigenous participants, samples were collected from 13 Yolngu individuals and 4 people from neighbouring NPN speakers and their nuclear genomes sequenced to a 30lil coverage. Using the short-read DNA BGISEQ-500 technology, these sequences were mapped to a reference genome and identified [~]24.86 million Single Nucleotide Variants (SNVs). The Yolngu SNVs were then compared to those of 36 individuals from 10 other Indigenous populations/locations across Australia and four worldwide populations using multidimensional scaling, population structure, F3 statistics and phylogenetic analyses. ResultsUsing the above methods, we infer that Yolngu speakers are closely related to neighbouring NPN speakers, followed by the Weipa population. No European or East Asian admixture was detected in the genomes of the Yolngu speakers studied here, which contrasts with the genomes of many other PN speakers that have been studied. Our results show that Yolngu speakers are more closely related to other PN speakers in the northeast of Australia than to those in central and western Australia studied here. Yolngu and the other Australian populations from this study share Papuans as an out-group. DiscussionThe study presented here provides an account of the nuclear and mitochondrial genomic diversity within the PN Yolngu Aboriginal population. The results show the Yolngu sample and their NPN neighbours have a strong genetic relationship. They also offer evidence of ancestral links between the Yolngu and PN-speaking populations in Cape York. From earlier fingerprint studies, consistent with the genomic results shown here, we suggest that there was a movement of people from the east into northeast Arnhem Land, associated with the flooding of the Sahul Shelf, and that this occurred between about 11 Kya and 8 Kya ago. Several Yolngu myths point to such a movement. It is suggested that the spread of the PN language or its speakers may have influenced the population structure of the Yolngu. Further genomic studies, with larger samples, of populations to the east of the Yolngu around the Gulf of Carpentaria into Cape York are required to test this hypothesis. Our results imply that PN did not spread with the movement of people across the continent, rather, the PN languages diffused among the different populations. It seems clear that the languages dispersed and not the people. The low level of relatedness detected between the Yolngu people and the people of the central arid desert of Australia suggests a long period of separation with different patterns of migration. Beyond Australia, Yolngu are most closely related to the Papuan people of New Guinea.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

White, N., Kumar, M., Lambert, D.. 2025-04-03. A high-resolution genomic study of the Pama-Nyungan speaking Yolngu people of northeast Arnhem Land, Australia.. https://doi.org/10.1101/2025.04.03.646226

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗

Large language model-based bibliometric evaluation of population descriptors in human genetics

As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.

genetics↗

Mitigating biases of rescaling in forward-in-time population genetic simulations

Forward-in-time population genetic simulations are widely used in evolutionary analyses, but simulating large populations and long genomic regions remains computationally demanding. To reduce this cost, parameter rescaling is widely employed, in which the original evolutionary process is approximated by one with a smaller population size and fewer generations. Recently, several studies using the SLiM simulator have raised concerns about the accuracy of this rescaling approach. In this study, we show that many of the biases reported in these studies can be mitigated by using a different simulation algorithm. These results reveal that the accuracy of parameter rescaling depends on how well the simulation algorithm preserves diffusion-limit properties under rescaling.

genetics↗