bioRxiv Science⌕ Search

Biology subjects

Stepec, D.

Publications and source records attributed to Stepec, D..

5 recordsLinked to original sources

Integrating Narrow-Window DIA with AI-Powered Search Enables Deep and Reliable Functional Proteomics

Narrow-window data-independent acquisition (nDIA) is emerging as a powerful technique for bottom-up proteomics. Here, we systematically benchmarked nDIA, wide-window DIA (wDIA), narrow-window data-dependent acquisition (nDDA), and wide-window DDA (wDDA) for rapid, single-shot proteomic analysis. For data processing, we introduced Tesorai Search, a new search engine leveraging a large pre-trained model and compared it with DIA-NN and FragPipe across both DIA and DDA datasets. Among 12 acquisition-analysis pipelines evaluated, nDIA combined with DIA-NN and Tesorai Search delivered the highest proteome coverage, identifying 10,255 and 10,766 protein groups from benchmark samples, respectively. Both search engines maintained rigorous false-discovery rate (FDR) control. While nDIA generally outperformed nDDA in sensitivity, FragPipe-DDA+ approach proved to be the most sensitive within the nDDA pipelines. However, entrapment analyses indicate that this sensitivity comes at the cost of less robust FDR control compared to Tesorai Search. As a proof of concept, we applied nDIA-MS to 17 cancer cell lines harboring DNA damage response (DDR) gene knockouts, successfully detecting significant downregulation of all targeted proteins and uncovering 81 DDR-related proteins modulated in at least one cell line. These results underscore nDIA-MS, together with DIA-NN and Tesorai Search, as a robust and scalable platform for high-throughput functional proteomic screening.

systems biology↗

Rapid Mechanistic Bridging of an Alzheimer's Disease Plasma Protein Staging Panel Across Brain Proteomics Cohorts

Background: Blood-based biomarkers are transforming Alzheimer's Disease (AD) diagnosis and staging. Recent multi-protein plasma panels accurately identify individuals with advanced Braak pathology, but whether these circulating biomarkers truly reflect the molecular remodeling underlying AD neuropathology remains unclear. Pre-existing datasets could answer such questions, but their reuse requires harmonized protein expression data, standardized sample annotations, and consistent study metadata. Methods: A recently reported seven-protein plasma staging panel was evaluated across multiple post-mortem proteomics brain cohorts using pre-structured datasets on the Tesorai platform. Five datasets with Braak stage data were identified and re-analyzed. A linear model was used to distinguish late (V-VI) from early Braak stage (0-IV). Because phosphorylated tau 217 (p-tau217) and amyloid beta 1-40 (A{beta}40) were not reported for any studies, raw spectra were reprocessed to quantify these proteoforms. Results: Across the cohorts, the plasma-derived biomarker panel consistently discriminated early from late disease despite heterogeneous protein coverage. Reprocessing of raw spectra recovered tau phosphopeptides absent from the original protein summaries, enabling inclusion of p-tau217, while A{beta}40 remained undetected. Together, these findings show that proteins comprising a recently proposed blood-based staging panel are associated with proteomic remodeling in the AD brain across independent cohorts. Conclusions: These findings provide biological validation for a recently proposed blood-based seven-protein staging panel by demonstrating that its constituent biomarkers are associated with disease-stage proteomic changes in AD brain tissue. More broadly, re-mining legacy mass spectrometry data for disease-relevant proteoforms, combined with pre-structured datasets, can accelerate evaluation of emerging blood-based biomarkers against neuropathological and molecular features of disease.

neuroscience↗

Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.

bioinformatics↗

Functional diversification of the cephalopod proteome by RNA-editing

Coleoid cephalopods exhibit the highest levels of ADAR-mediated RNA editing of any known animal, yet the functional consequences of most recoding events remain largely unknown. We integrate proteomics with biochemical and cellular assays to characterize thousands of recoding events across the Doryteuthis pealeii proteome. Using quantitative and functional mass spectrometry, we show that RNA edit-driven recoding reshapes the cellular proteome to alter protein stability, subcellular localization, post-translational modifications, and enzymatic activity. --Recoding can regulate post-translational modifications through their creation or ablation, and this has direct effects on protein function and protein-protein interactions. Recoding of the E3 ligase MARCHF5 drives widespread changes in substrate ubiquitylation and perturbs mitochondrial homeostasis, illustrating how RNA editing can influence organelle function. These data provide the first proteome-scale view of how extensive RNA recoding diversifies protein function in coleoid cephalopods and offers a new framework for understanding how RNA-level plasticity shapes protein function and cellular physiology.

molecular biology↗

Tesorai Search: Large pretrained model boosts identifications in mass spectrometry proteomics without the need for Percolator.

The original mass spectrometry search engines used simple algorithms for peptide identification. Recent tools improved accuracy by adding several extra components such as fragment ion intensities or retention times prediction and training target-decoy classifiers on-the-fly, leading to sometimes inconsistent results. Our study explores the impact of replacing those extra components with a deep-learning pretrained model that directly learns the complex relationship between the full spectra and associated peptide sequence, without using decoys. This simplified workflow has fewer parameters to tweak, making it easier to use and perform robustly on data from instruments and use-cases never seen during training. Surprisingly, our approach consistently identifies more peptides than FragPipe, PEAKS, and Proteome Discoverer (12%, 9%, and 21% more, respectively, across a range of datasets). Tesorai Search is also fast - 250 immunopeptidomics searches in 45 minutes - and free for academics, available as a webserver at console.tesorai.com.

bioinformatics↗