bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.09.28.755120

LOAM: A family of genomic language models trained on long-read soil metagenomes

Abstract

Soil ecosystems represent a vast, largely uncharacterised reservoir of microbial diversity. Metagenomic assembly has unlocked access to this resource, and recent advances in long-read sequencing have improved the recovery and quality of microbial genomes from complex samples. While genomic language models have proven highly effective at capturing biological concepts from large sequence datasets, they have predominantly been trained on reference genomes or short-read assemblies. Here, we present LOAM, a family of decoder-only genomic language models ranging from 25 to 624 million parameters and trained exclusively on Oxford Nanopore long-read environmental metagenomes. LOAM model performance scales predictably with model size and training-token budget. Despite a relatively small training sequence corpus, LOAM models outperformed comparably sized models across biological benchmarks, and achieved performance competitive with substantially larger state-of-the-art models trained on much larger datasets. Context-intervention experiments further showed that LOAM models use genomic information over several kilobases, highlighting the potential value of increased contiguity provided by long-read metagenome-assembled genomes. For probe-based benchmark tasks, we systematically evaluated representations across hidden layers and revealed that task-relevant biological information was frequently more linearly accessible from intermediate than final model layers. Finally, we observed that variation in zero-shot variant-effect prediction was strongly associated with the presence of homologous target sequences in the pre-training corpus. Together, these results establish long-read environmental metagenomes as a viable foundation for training competitive genomic language models and demonstrate the importance of both model scale and training-corpus composition for biological generalisation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ferry, Q. R., Frank, M., Jehn, J., Schaks, M., Steinkraus, B. R., Rajakumar, T.. 2026-10-02. LOAM: A family of genomic language models trained on long-read soil metagenomes. https://doi.org/10.64898/2026.09.28.755120

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

X-Y divergence of house fly (Musca domestica) proto-sex chromosomes follows distinct evolutionary trajectories despite residing in the same genome

Sex determination systems and sex chromosomes frequently differ between species. One cause of these differences is new sex determining genes that drive evolutionary turnover of sex chromosomes. At the earliest stages of this turnover, X and Y (or Z and W) chromosomes start out as nearly identical homologs, and they can diverge via chromosomal rearrangements (e.g., inversions) that suppress X-Y recombination. However, multiple examples from across animals and plants provide exceptions to this canonical model of sex chromosome evolution. For example, some X-Y pairs remain undifferentiated for long evolutionary time periods, while other sex chromosomes become differentiated without chromosomal rearrangements. How or why these non-canonical trajectories occur remains elusive, despite increasing evidence of their pervasiveness. The house fly, Musca domestica, is a well-suited system to address this gap because all six chromosomes can be a Y, providing multiple replicates of a natural experiment within a single genomic environment. To test for canonical and non-canonical evolutionary trajectories, we generated haplotype-resolved chromosome-level assemblies from five strains of the house fly, each of which carries a different Y chromosome (IM, IIM, IIIM, VM, and YM). We identified an inversion on only one of the sex chromosomes (IIM), which was associated with elevated X-Y divergence but did not capture the male-determining locus. In contrast, there was X-Y divergence across almost the entire length of the IM, IIIM, and VM sex chromosomes, despite no detectable inversions. YM was the only sex chromosome to contain substantial Y-specific sequences, which were limited to a segment on one end of the chromosome containing the male-determining gene. This YM chromosome, and its corresponding X, were highly diverged from the X chromosome found in many other flies (Muller element F), despite a strong cytological resemblance. This study highlights how multiple different canonical and non-canonical modes of sex chromosome evolution can co-exist within a single genome.

genomics↗

Method-dependent biases in cell type detection between single-cell and single-nucleus RNA sequencing in the photosymbiotic acoel Praesagittifera naikaiensis

Background Comparisons of single-cell and single-nucleus RNA sequencing (scRNA-seq and snRNA-seq) data have been described in some mammalian tissues and, subsequently, in Drosophila, but remain unexplored in most invertebrate lineages. The xenacoelomorphs occupy key phylogenetic positions, yet they differ anatomically from mammals. They have a reduced extracellular matrix, high-salt body fluid, and no circulatory system. Despite these differences, they possess a well-developed nervous system. One of the xenacoelomorphs, the photosymbiotic acoel (Praesagittifera naikaiensis) also harbours symbiotic Tetraselmis algae, whose RNA can be co-captured with host RNA. Results We compared scRNA-seq and snRNA-seq data from whole P. naikaiensis specimens. Both methods yielded high-quality data with comparable gene detection but a larger share of scRNA-seq reads derived from symbiotic algae. Gene-level analyses revealed that neural genes were enriched in snRNA-seq relative to non-neural genes. Cross-method label transfer and integration-based validation identified six snRNA-seq clusters, including some neural populations, that lacked a clear scRNA-seq counterpart. In contrast, three cell populations, including muscle and metabolically active clusters, were reciprocally validated as captured by both methods. Conclusions Our results show that key snRNA-seq advantages, particularly the enhanced recovery of neural transcripts, are recapitulated in our dataset, which is consistent with previous reports in mammals. Furthermore, snRNA-seq reduces symbiont-derived reads and recovers several cell populations underrepresented in scRNA-seq. These findings provide practical guidance for cell atlas construction in non-model, symbiotic invertebrates.

genomics↗

ST-DISTAL: Dual-Branch Graph Convolution with Distributional Alignment for Cell-Type Deconvolution

Spatial transcriptomics enables high-throughput gene expression profiling while preserving spatial information, offering valuable insights into tissue architecture and cellular organization. Methods that achieve single-cell or subcellular spatial resolution typically rely on predefined gene panels, which limits genome-wide discovery. In contrast, sequencing-based spatial transcriptomics platforms provide broad transcriptome coverage but measure expression at lower spatial resolution, capturing mixtures of multiple cell types within each spatial location and thereby requiring accurate cell-type deconvolution. Many existing deconvolution methods do not jointly model molecular expression similarity and spatial context, fail to align predicted cell-type compositions with biological priors, and may produce spatially inconsistent results that do not reflect the smooth and structured organization of real tissues. We present ST-DISTAL, a dual-branch graph convolutional framework for cell-type deconvolution in spatial transcriptomics data. ST-DISTAL integrates conventional spatial graph convolution with spectral Chebyshev filtering through an adaptive attention mechanism, enabling the model to capture both local spatial dependencies and global structural patterns. The framework is trained using a composite objective that jointly enforces accurate cell-type proportion estimation, global distributional alignment with reference profiles, and spatial smoothness. Evaluations on three simulated benchmark datasets demonstrate that ST-DISTAL consistently outperforms state-of-the-art methods in both prediction accuracy and distributional alignment. Further validation on a real human embryonic heart dataset and a colorectal cancer dataset shows biologically plausible and spatially coherent cell-type organization.

genomics↗