bioRxiv · 10.1101/2022.11.03.515084
Adapting protein language models for rapid DTI prediction
Abstract
We consider the problem of sequence-based drug-target interaction (DTI) prediction, showing that a straightforward deep learning architecture that leverages pre-trained protein language models (PLMs) for protein embedding outperforms state of the art approaches, achieving higher accuracy, expanded generalizability, and an order of magnitude faster training. PLM embeddings are found to contain general information that is especially useful in few-shot (small training data set) and zero-shot instances (unseen proteins or drugs). Additionally, the PLM embeddings can be augmented with features tuned by task-specific pre-training, and we find that these task-specific features are more informative than baseline PLM features. We anticipate such transfer learning approaches will facilitate rapid prototyping of DTI models, especially in low-N scenarios.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sledzieski, S., Singh, R., Cowen, L., Berger, B.. 2022-11-04. Adapting protein language models for rapid DTI prediction. https://doi.org/10.1101/2022.11.03.515084
Cite the original work for its findings. Save a collection to share your selection of sources.