bioRxiv · 10.1101/2023.05.31.542682
genomicBERT and data-free deep-learning model evaluation
Abstract
The genome, which serves as the inherent language directing the blueprint of life, offers significant analysis prospects by combining Natural Language Processing (NLP) and machine learning (ML). Integrating biological sequences with other digital healthcare information has potential to transform data-driven diagnostics. Large language models (LLMs) can be harnessed to decode the genomic language. This endeavor encounters three critical challenges: First, long biomolecular sequences require segmentation into smaller subunits, which is non-trivial since many biological "words" remain unknown. Second, the analysis of extended DNA sequences using LLMs demands a compute-intensive infrastructure. Third, ensuring reproducibility and reusability of modeling workflows remains an unresolved issue. To tackle these challenges, we introduce an empirical DNA tokenisation approach and a versatile, semantic-aware, genome language model --genomicBERT. The model is species-agnostic and operates seamlessly at the DNA or RNA levels. By introducing a reduced and specialized DNA vocabulary, our approach minimizes computational overhead and optimizes performance. Our benchmarking demonstrates that the genomicBERT matches or surpasses the performance of contemporary tools on the same datasets under different experimental conditions. To encourage collaboration and ease of access, we introduce genomicBERT as an integral component of the openly accessible conda package, genomeNLP. Validated across diverse case studies, genomicBERT lowers the barriers to decoding genomic language, relying solely on sequence data to extract meaningful insights. HighlightsO_LIThis novel model offers a compelling solution for DNA sequence analysis by significantly reducing model size and computational costs without compromising performance, setting a new standard for efficient model development. C_LIO_LIWe demonstrate that a powerful vocabulary and tokenization method helps to derive patterns from biological sequence data while accounting for hidden semantic rules. C_LIO_LIOur method is agnostic to species or biomolecule type as it is data-driven. Hence, it can be applied to DNA and RNA C_LIO_LIWe validate the important genomicBERT tokens by mapping back to the biologically significant motifs. C_LIO_LIWe present a publicly available genome language modeling toolkit called genomeNLP, specifically designed to combine computational linguistics and genomics, enabling researchers from biology backgrounds to analyze and interpret genomic sequences effectively. C_LI
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, T., Tyagi, N., Chauhan, S., Peleg, A. Y., Tyagi, S.. 2023-06-01. genomicBERT and data-free deep-learning model evaluation. https://doi.org/10.1101/2023.05.31.542682
Cite the original work for its findings. Save a collection to share your selection of sources.