bioRxiv · 10.1101/2023.07.11.548628
DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks
Abstract
Pre-trained large language models demonstrate potential in extracting information from DNA sequences, yet adapting to a variety of tasks and data modalities remains a challenge. To address this, we propose DNAGPT, a generalized DNA pre-training model trained on over 200 billion base pairs from all mammals. By enhancing the classic GPT model with a binary classification task (DNA sequence order), a numerical regression task (guanine-cytosine content prediction), and a comprehensive token language, DNAGPT can handle versatile DNA analysis tasks while processing both sequence and numerical data. Our evaluation of genomic signal and region recognition, mRNA abundance regression, and artificial genome generation tasks demonstrates DNAGPTs superior performance compared to existing models designed for specific downstream tasks, benefiting from pre-training using the newly designed model structure.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhang, D., Zhang, w., He, B., Zhang, J., Qin, C., Yao, J.. 2023-07-12. DNAGPT: A Generalized Pretrained Tool for Multiple DNA Sequence Analysis Tasks. https://doi.org/10.1101/2023.07.11.548628
Cite the original work for its findings. Save a collection to share your selection of sources.