bioRxiv · 10.1101/2024.05.30.596740
ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning
Abstract
ProTrek redefines protein exploration by seamlessly fusing sequence, structure, and natural language function (SSF) into an advanced tri-modal language model. Through contrastive learning, ProTrek bridges the gap between protein data and human understanding, enabling lightning-fast searches across nine SSF pairwise modality combinations. Trained on vastly larger datasets, ProTrek demonstrates quantum leaps in performance: (1) Elevating protein sequence-function interconversion by 30-60 fold; (2) Surpassing current alignment tools (i.e., Foldseek and MMseqs2) in both speed (100-fold acceleration) and accuracy, identifying functionally similar proteins with diverse structures; and (3) Outperforming ESM-2 in 9 of 11 downstream prediction tasks, setting new benchmarks in protein intelligence. These results suggest that ProTrek will become a core tool for protein searching, understanding, and analysis.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Su, J., Zhou, X., Zhang, X., Yuan, F.. 2024-06-03. ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning. https://doi.org/10.1101/2024.05.30.596740
Cite the original work for its findings. Save a collection to share your selection of sources.