bioRxiv · 10.64898/2026.09.10.750553
A foundation model learns the sequence and functional grammar of fully human heavy-chain-only antibodies
Abstract
Antibody language models learn from large natural repertoires, but whether generic representations capture the constraints of specialized antibody formats remains unclear. We first characterized fully human heavy-chain-only antibodies (HCAbs) independently of HCAb-trained models. Source-aware comparisons with conventional human VH domains revealed a reproducible distributional shift localized predominantly to CDR1/2, CDR3 architecture and, where supported, a restricted framework region rather than widespread framework remodeling. These model-independent differences motivated repertoire-specific pretraining. We developed HCAbLM, to our knowledge the first foundation model pretrained specifically on a large-scale fully human HCAb repertoire, using 31.8 million sequences from 73 independently immunized HCAb mice. HCAbLM learned a region-selective sequence-compatibility prior distinct from conventional antibody language models, and its frozen representations transferred to experimentally measured SEC purity, HIC behavior and thermal stability in grouped internal validation and retrospective cross-project evaluation. These findings identify repertoire composition as an important biological design variable for foundation models of specialized antibody formats.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nona Biosciences AI4S Team,, Miao, H.. 2026-09-15. A foundation model learns the sequence and functional grammar of fully human heavy-chain-only antibodies. https://doi.org/10.64898/2026.09.10.750553
Cite the original work for its findings. Save a collection to share your selection of sources.