bioRxiv · 10.1101/2024.07.02.601703
BiRNA-BERT Allows Efficient RNA Language Modeling with Adaptive Tokenization
Abstract
Recent advancements in Transformer-based models have spurred interest in their use for biological sequence analysis. However, adapting models like BERT is challenging due to sequence length, often requiring truncation for proteomics and genomics tasks. Additionally, advanced tokenization and relative positional encoding techniques for long contexts in NLP are often not directly transferable to DNA/RNA sequences, which require nucleotide or character-level encodings for tasks such as 3D torsion angle prediction. To tackle these challenges, we propose an adaptive dual tokenization scheme for bioinformatics that utilizes both nucleotide-level (NUC) and efficient BPE tokenizations. Building on the dual tokenization, we introduce BiRNA-BERT, a 117M parameter Transformer encoder pretrained with our proposed tokenization on 28 billion nucleotides across 36 million coding and non-coding RNA sequences. The learned representation by BiRNA-BERT generalizes across a range of applications and achieves state-of-the-art results in long-sequence downstream tasks and achieves a performance comparable to 6x larger models in short-sequence tasks with 27xless pre-training compute. BiRNA-BERT can dynamically adjust its tokenization strategy based on sequence lengths, utilizing NUC for shorter sequences and switching to BPE for longer ones, thereby offering, for the first time, the capability to efficiently handle arbitrarily long DNA/RNA sequences. 1
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tahmid, M. T., Shahgir, H. S., Mahbub, S., Dong, Y., Bayzid, M. S.. 2024-07-04. BiRNA-BERT Allows Efficient RNA Language Modeling with Adaptive Tokenization. https://doi.org/10.1101/2024.07.02.601703
Cite the original work for its findings. Save a collection to share your selection of sources.