bioRxiv Science⌕ Search

Biology subjects

A, J.

Publications and source records attributed to A, J..

5 recordsLinked to original sources

MassNet: billion-scale AI-friendly mass spectral corpus enables robust de novo peptide sequencing

Breakthroughs in artificial intelligence (AI) for natural language processing and computer vision have been largely driven by high-quality, large-scale datasets such as OpenWebText and ImageNet. Inspired by this, we present MassNet, a foundational resource for proteomics designed to accelerate deep learning applications. MassNet is the largest known corpus of data-dependent acquisition (DDA) mass spectrometry (MS) data, derived from ~30 TB of raw files and comprising 1.54 billion MS/MS spectra, resulting in 558 million peptide-spectrum matches (PSMs) across 35 species, including animals, plants, and microbes. Within the human subset, MassNet includes more than 1.7 million precursors and 19,966 proteins, covering 98% of annotated human proteins. To enable efficient AI training, we developed the Mass Spectrometry Data Tensor (MSDT), a structured format based on Parquet that enables standardized, high-performance batch access and seamless integration with GPU and TPU platforms for distributed training. We further extended MassNet to support de novo peptide sequencing, which infers peptide sequences directly from MS/MS spectra without reference databases, and is critical for discovering novel proteins, characterizing non-model organisms, and identifying post-translational modifications (PTMs). We introduce XuanjiNovo, a non-autoregressive Transformer model that leverages a curriculum learning strategy to enhance training stability. By dynamically adjusting learning difficulty based on model performance, XuanjiNovo achieves smooth convergence on complex, multi-distributional data without manual hyperparameter tuning. Trained on 100 million PSMs from the MassNet, it consistently outperforms state-of-the-art methods across diverse benchmarking tasks. Peptide recall exceeds 0.8 on the Bacteroides thetaiotaomicron and Zea mays datasets. On human data acquired using the Orbitrap Astral platform, XuanjiNovo achieves achieves 38.8% to 144.3% improvement over existing models. MassNet represents the first large-scale, standardized foundational dataset in proteomics, marking a critical milestone in the integration of artificial intelligence into proteomics research.

bioinformatics↗

Spatial distribution of the proteome in human body and cancers

A comprehensive spatial distribution of the proteome in human body and cancers is fundamental for understanding human biology and diseases including cancers. Here, we present an anatomically resolved human proteome derived from 1781 benign and malignant samples from 58 major tissue types encompassing 251 specific tissues and 25 carcinomas. Based on a spectral library covering over 75% of the human protein-coding genes with 208 understudied and 82 missing proteins characterized, we quantified over 13,000 proteins in these samples using data-independent acquisition proteomics. This data resource presents the so far most comprehensive quantitative proteomic landscape of human tissues and common carcinomas. It allows systematic evaluation of tissue-specific drug responses, identification of drug candidates that may be repurposed as antineoplastics, and discovery of novel targets for anticancer therapy. This resource, available as an online knowledgebase, refines our knowledge of spatial distribution of the human proteome and tumor-specific protein modulation.

systems biology↗

DyNDG: Identifying Leukemia-Related Genes based on the Time-Series Dynamical Network by Integrating Differential Genes

Leukemia is a malignant disease of progressive accumulation characterized by high morbidity and mortality rates, and investigating its disease genes is crucial for understanding its etiology and pathogenesis. Network propagation methods have emerged and been widely employed in disease gene prediction, but most of them focus on static biological networks, which hinders their applicability and effectiveness in the study of progressive diseases. Moreover, there is currently a lack of special algorithms for the identification of leukemia disease genes. Here, we proposed DyNDG, a novel Dynamic Network-based model, which integrates Differentially Expressed Genes to identify leukemia-related genes. Initially, we constructed a time-series dynamic network to model the development trajectory of leukemia. Then, we built a background-temporal multilayer network by integrating both the dynamic network and the static background network, which was initialized with differentially expressed genes at each stage. To quantify the associations between genes and leukemia, we extended a random walk process to the background-temporal multilayer network. The experimental results demonstrate that DyNDG achieves superior accuracy compared to several state-of-the-art methods. Moreover, after excluding housekeeping genes, DyNDG yields a set of promising candidate genes associated with leukemia progression or potential biomarkers, indicating the value of dynamic network information in identifying leukemia-related genes. The implementation of DyNDG is available at https://github.com/CSUBioGroup/DyNDG.

bioinformatics↗

DDA-BERT: leveraging transformer architecture pre-training for data-dependent acquisition mass spectrometry-based proteomics

In data-dependent acquisition mass spectrometry (DDA-MS)-based proteomics, machine learning based rescoring methods are often employed to integrate multiple scores measuring quality of peptide-spectrum matches (PSMs) from different aspects. Existing rescoring tools face limitations incurred by manual feature extraction, shallow machine learning models, and limited training data. Here, we introduce DDA-BERT, a transformer-based end-to-end deep learning model trained with 95 million human spectra for PSM rescoring. DDA-BERT demonstrates superior performance across various sample types, mass spectrometer platforms, trace sample proteomics, and multiple species proteome data. It consistently outperforms state-of-the-art methods of its kind for DDA-MS spectra analysis with up to 103.5% increased protein group identifications. For single cell proteomics, DDA-BERT identifies up to 85.6% more protein groups compared to existing tools. In addition, DDA-BERT can be effectively extended to proteome of non-human species. DDA-BERT offers a robust and scalable solution for enhancing PSM rescoring in proteomics.

bioinformatics↗

DPHL v2: An updated and comprehensive DIA pan-human assay library for quantifying more than 14,000 proteins.

A comprehensive pan-human spectral library is critical for biomarker discovery using mass spectrometry (MS)-based proteomics. DPHL v1, a previous pan-human library built from 1096 data-dependent acquisition (DDA) MS data of 16 human tissue types, allows quantifying 10,943 proteins. However, a major limitation of DPHL v1 is the lack of semi-tryptic peptides and protein isoforms, which are abundant in clinical specimens. Here, we generated DPHL v2 from 1608 DDA-MS data acquired using Orbitrap mass spectrometers. The data included 586 DDA-MS newly acquired from 17 tissue types, while 1022 files were derived from DPHL v1. DPHL v2 thus comprises data from 24 sample types, including several cancer types (lung, breast, kidney, and prostate cancer, among others). We generated four variants of DPHL v2 to include semi-tryptic peptides and protein isoforms. DPHL v2 was then applied to a publicly available colorectal cancer dataset with 286 DIA-MS files. The numbers of identified and significantly dysregulated proteins increased by at least 21.7% and 14.2%, respectively, compared with DPHL v1. Our findings show that the increased human proteome coverage of DPHL v2 provides larger pools of potential protein biomarkers.

bioinformatics↗