bioRxiv · 10.1101/2022.05.19.492714
Convolutions are competitive with transformers for protein sequence pretraining
Abstract
Pretrained protein sequence language models have been shown to improve the performance of many prediction tasks, and are now routinely integrated into bioinformatics tools. However, these models largely rely on the Transformer architecture, which scales quadratically with sequence length in both run-time and memory. Therefore, state-of-the-art models have limitations on sequence length. To address this limitation, we investigated if convolutional neural network (CNN) architectures, which scale linearly with sequence length, could be as effective as transformers in protein language models. With masked language model pretraining, CNNs are competitive to and occasionally superior to Transformers across downstream applications while maintaining strong performance on sequences longer than those allowed in the current state-of-the-art Transformer models. Our work suggests that computational efficiency can be improved without sacrificing performance simply by using a CNN architecture instead of a Transformer, and emphasizes the importance of disentangling pretraining task and model architecture.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yang, K. K., Lu, A. X., Fusi, N. K.. 2022-05-20. Convolutions are competitive with transformers for protein sequence pretraining. https://doi.org/10.1101/2022.05.19.492714
Cite the original work for its findings. Save a collection to share your selection of sources.