bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.06.17.660172

Seq2KING: An unsupervised internal transformer representation of global human heritages

Abstract

Determining the intricate tapestry of human genetic relationships is a central challenge in population genetics and precision medicine. We propose that the principles of lexical connectivity, which words derive meaning from their contextual interactions, can be adapted to genetic data, enabling transformer models to reveal that individuals with higher genetic similarity form stronger latent connections. We explored this by transposing KING kinship-related matrices into the (query, key, value) QKV latent space within transformer models and determined that attention mechanisms can capture genetic relatedness in an unsupervised fashion. We found that individuals had an attention weight connectivity of 85.34% (p<0.05) if they were from within the same continent, compared to if they were from other continents. Surprisingly, we found that some encoder layers required inversion of their latent representations for this connectivity to become obvious. Lastly, we used BERTViz to create human-readable hyper-dense connectivity patterns among individuals. Our approach is purely based on attention, which yields a non-discrete spectrum of relatedness, and thus uncovers patterns on first principles. Seq2KING addresses the significant challenge of discovering population structures to construct a global human relatedness map, without relying on predefined labels. Our excavation into the latent space is a paradigm shift from legacy-supervised genetic methodologies, which presents a new way to understand the human pangenome as well as discern population substructures for creating precision genetic medicines. Non-Expert DescriptionIs it possible to build artificial intelligence (AI) to read the human genome as a first language? Why would one want such AI? We at Ecotone believe that such AI will provide the genetic coordinates needed to manufacture CRISPRs medications to cure about [~]10,000 genetic diseases. How does one build such AI? Our recently released model dnaSORA proposed a means to assign meaning to every single token (typically referred to as a base) of all 3 billion tokens in the human genome (Koreniuk & Njie, 2025). This builds the vocabulary for reading the human genome as a first language. For dnaSORA to work, it needs to know the heritages of people that are in its model of our genetics. We mostly rely on country, culture and geography to determine our heritages, but this is too error-prone for dnaSORA. Also error-prone in our experience are legacy genetic approaches such as those used by 23andMe. Our research here introduces Seq2KING, a new artificial intelligence method that is based on excavating the insides of transformers to uncover hidden patterns of genetic relatedness among people around the world--without needing any prior labels or categorizations. The key innovation of Seq2KING is applying the principles of lexical connectivity-- how words derive meaning through their relationships to other words--to genetic data. Just as "dog" gains meaning through its connections to words like "pet," "animal," and "loyal," we show that individuals genomes can be understood through their genetic connections to others. We start by converting raw genetic data into a compact kinship matrix (using a tool called KING) that summarizes how closely everyone is related. We then feed these kinship values into a transformer model--the same kind of AI behind cutting-edge language tools like ChatGPT. Inside the transformer, special components called "attention heads" learn which individuals are most similar, strengthening links between people from the same region and showing subtler connections across continents. Unlike legacy approaches that rely on discrete pre-defined categories, Seq2KING provides continuous measures of relatedness, allowing us to visualize connections between any individual and all other humans. Additionally, because Seq2KING operates directly within the transformers internal reasoning system, it can be seamlessly integrated as a component within larger genome interpretation systems--essentially functioning like high-speed cache memory for heritage assignments, dramatically improving both efficiency and scalability. By examining these attention patterns, we can reconstruct familiar population groupings--such as European, African, and Asian heritage--entirely by the models internal logic. Finally, we use a visualization technique (BERTViz) to turn these dense connection maps into intuitive diagrams that highlight population connections between individuals. Because our approach doesnt rely on pre-assigned labels, it offers a truly unbiased way to explore human population structure. This could help scientists trace migration routes that resulted in the peopling of the continents, find subtle subgroups within larger populations, and remove "background noise" in genetic studies of disease. Ultimately, Seq2KING paves the way for more precise genetic maps of all humans, revealing the natural "family trees" hidden in our DNA and bringing us one step closer to reading the human genome as a first language.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jonnalagadda, B., Njie, e.. 2025-06-23. Seq2KING: An unsupervised internal transformer representation of global human heritages. https://doi.org/10.1101/2025.06.17.660172

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

The histone demethylase Kdm5 and the ARGONAUTE proteins Piwi and Aubergine regulate female abdominal pigmentation in Drosophila melanogaster

Insect pigmentation is an ecologically critical trait influencing many physiological processes. In Drosophila melanogaster, abdominal pigmentation is sexually dimorphic: males have fully pigmented posterior segments, while females exhibit a posterior melanin stripe. Pigmentation relies on the expression of pigmentation genes that encode enzymes involved in pigment synthesis. These genes are tightly regulated during pupal and young adult stages. To expand the gene regulatory network of pigmentation genes, we conducted an RNAi screen using the yellow-Gal4 driver, expressed during the pupal stage in abdominal epidermis. One of the candidates from this screen, Kdm5, encodes a histone demethylase erasing the H3K4me3 histone mark catalyzed by the histone methyl-transferase Trithorax (Trx). We show that Kdm5 down-regulation reduces abdominal pigmentation, mimicking trx down-regulation. Kdm5 activates melanin production through regulation of the pigmentation gene tan. Transcriptomic analyses reveal that Kdm5 and Trx share many targets in pupal abdominal epidermis, including piRNA pathway components such as piwi and aubergine. These piRNA components, originally associated with transposon silencing in the germline, also function in some somatic tissues such as the nervous system, the fat body or the gut. We demonstrate that Piwi and Aubergine participate in female abdominal pigmentation establishment, without evident piRNA production. We also show that Kdm5 and Piwi act not only in pupal abdominal epidermis but also in pupal fat body. This study therefore expands the regulatory network of pigmentation genes. It identifies a new somatic function for Kdm5 and Piwi and reveals a role for pupal fat body in female abdominal pigmentation regulation.

genetics↗

Genetic diversity within and between polyploid sugarcane (Saccharum spp.) families obtained via caryopsis using microsatellite markers and multicategory model

Genetic diversity analyses are essential for sugarcane (Saccharum spp.) breeding programs. Crossbreeding, based on genetic distances between parental plants, is a tool used to increase genetic variability and enhance plant selection; however, quantifying variation in highly polyploid species remains a challenge. The present study aimed to evaluate the diversity within and between 12 families of sugarcane derived from caryopses, analyzing 120 individual seedlings arranged in an augmented block design. Genotyping was performed using primers for 16 microsatellite loci, five simple sequence repeat (SSR) loci, and 11 expressed sequence tag-SSR (EST-SSR) loci. To accurately account for polyploidy, similarity calculations were performed using Bruvos distances among individuals and RST distances among the families. Analysis of molecular variance (AMOVA) indicated that most of the genetic variability was within families (72%), with only 28% found between them. This high level of intra-family variation demonstrates that a significant reservoir of genetic diversity remains available within the crosses. The highest genetic similarity was observed between the families RB986952 x RB986960 and RB036122 x RB03611, whereas the lowest genetic similarity was observed between the families RB97319 x RB966928 and RB106802 x RB855036. Although the evaluated families shared high genetic similarity, the pronounced genetic variation within them demonstrates a robust recombination potential, indicating that the genetic basis of sugarcane can be better explored using the high variability that already exists in the selection of desirable morpho-agronomic characteristics within the families. Furthermore, this study highlights the importance of using appropriate distances for diversity studies with codominant markers, such as microsatellites, in polyploid species.

genetics↗

Optimizing DNA extraction from environmentally degraded bone samples for molecular identification of cetacean species

Molecular identification of cetacean bone remains can be limited by DNA degradation and the presence of PCR inhibitors. Here, we present an optimized DNA extraction protocol based on a total demineralization method for environmentally exposed cetacean bones. The protocol uses 100 mg of bone powder, 24 h digestion with EDTA, N-lauroylsarcosine, and proteinase K, followed by a modified silica-column purification. Nine environmentally degraded bone samples representing eight individuals were processed. DNA concentrations ranged from 7.3 to 57.1 ng/uL (mean SD = 25.91- 13.91 ng/uL). The mitochondrial cytochrome b gene was successfully amplified from all samples using conventional PCR, and five samples (55.6%) yielded sequences suitable for downstream analysis. BLASTn identified Balaenoptera physalus as the closest database match for all recovered sequences, and phylogenetic analysis further supported their association with B. physalus reference sequences. These results demonstrate that the proposed protocol provides a practical approach for recovering amplifiable and molecularly informative mitochondrial DNA from environmentally degraded cetacean bone material, facilitating molecular identification from challenging skeletal remains.

genetics↗