bioRxiv Science⌕ Search

Biology subjects

Char, S.

Publications and source records attributed to Char, S..

3 recordsLinked to original sources

Protein language models and the long tail of functional diversity

Protein language model performance on downstream tasks depends on the pretraining data, motivating recent efforts to combine genomic- and metagenomic-derived protein sequences into large-scale atlases. Because these datasets are highly redundant, sequences are typically clustered by similarity and sampled during training. Sequences that do not belong to any cluster, known as "singletons", are typically excluded from training and evaluation because they are considered to be artifacts. However, singletons represent the long tail of functional diversity and are abundant in many large-scale atlases: nearly 43% of the 3.34 billion sequences in the joint genomic-metagenomic dataset GigaRef are singletons. Here, we characterize singletons derived from UniRef and GigaRef by assessing whether clustering missed homologs, how much their exclusion affects protein language model (PLM) training, and which biological domains they contain. We find that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations. We also show that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training. Finally, metagenomic singletons carry denser, more diverse domain content than clustered sequences, including domain-level homology that sequence-identity clustering misses. Together, these results support including singletons in PLM training and call for closer examination of data curation in large-scale integrated sequence atlases.

bioinformatics↗

The Dayhoff Atlas: scaling sequence diversity for improved protein generation

Organized information powers modern biology, a framework pioneered by Margaret Dayhoffs Atlas of Protein Sequence and Structure and advanced by todays databases and computational methods. Here, we extend this paradigm for the AI era, presenting the Dayhoff Atlas of protein sequence data and generative models to accelerate protein biology and design. The Atlas introduces GigaRef, the largest open dataset of natural proteins, spanning 3.34B genomic and metagenomic sequences across 1.70B clusters, and BackboneRef, which distills structural information from 240,811 synthetic backbones into 46M synthetic sequences. Leveraging these datasets, we trained the Dayhoff protein language models, which can predict mutation effects, scaffold structural motifs, and generate novel proteins within families. Training on metagenomic and structure-based synthetic sequences increased the expression rates of generated proteins, demonstrating the value of data diversity and scale. We release the Dayhoff Atlas code, datasets, and models under a permissive license to empower computation in protein design.

bioengineering↗

ProtNote: a multimodal method for protein-function annotation

Understanding the protein sequence-function relationship is essential for advancing protein biology and engineering. However, fewer than 1% of known protein sequences have human-verified functions. While deep learning methods have demonstrated promise for protein function prediction, current models are limited to predicting only those functions on which they were trained. Here, we introduce ProtNote, a multimodal deep learning model that leverages free-form text to enable both supervised and zero-shot protein function prediction. ProtNote not only maintains near state-of-the-art performance for annotations in its train set, but also generalizes to unseen and novel functions in zero-shot test settings. We envision that ProtNote will enhance protein function discovery by enabling scientists to use free text inputs, without restriction to predefined labels - a necessary capability for navigating the dynamic landscape of protein biology.

bioengineering↗