bioRxiv · 10.1101/2024.11.11.622930
DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis
Abstract
Data-independent acquisition mass spectrometry (DIA-MS) plays an increasingly important role in quantitative proteomics. Here, we introduce DIA-BERT, a software tool that leverages a transformer-based pre-trained artificial intelligence (AI) model for the analysis of DIA proteomics data. Over 276 million of high-quality peptide precursors extracted from existing DIA-MS files were used for training the identification model, while 34 million peptide precursors from synthetic DIA-MS files were used for training the quantification model. Compared to DIA-NN, DIA-BERT led to on average 54% more protein identifications and 37% more peptide precursors in five different human cancer (cervical cancer, pancreatic adenocarcinoma, myosarcoma, gallbladder cancer, and gastric carcinoma) sample sets with a high degree of quantitative accuracy. This study highlights the potential of utilizing pre-trained models and synthetic datasets to advance DIA proteomics analysis.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Liu, Z., Liu, P., Sun, Y., Nie, Z., Zhang, X., Zhang, Y., Chen, Y., Guo, T.. 2024-11-12. DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis. https://doi.org/10.1101/2024.11.11.622930
Cite the original work for its findings. Save a collection to share your selection of sources.