bioRxiv Science⌕ Search

Biology subjects

Kucharski, P.

Publications and source records attributed to Kucharski, P..

2 recordsLinked to original sources

Scaling and democratising structure-based protein function prediction with metagenomic-deepFRI

High-quality protein structure models have become widely available, offering insights into protein function, yet they remain underutilized. Here, we introduce metagenomic-deepFRI, a framework incorporating structural templates into functional annotation pipelines at speeds comparable to sequence-alignment methods. Notably, structural features improved GO term prediction confidence and Information Content by up to 50%. Applied to metagenomic datasets, the framework achieves nearly 90% annotation coverage, enabling protein function inference without explicit orthology-based transfer.

bioinformatics↗

Resolution of recursive data corruption to transform T-cell epitope discovery

Accurate prediction of MHC class I-presented peptides is essential for any vaccine or T-cell therapy design, yet reported gains on in silico benchmarks have not translated into clinical successes. Here we show that this discrepancy may come from a common methodological error: immunopeptidomics datasets are fundamentally contaminated by existing prediction models through prediction-based deconvolution and filtering, resulting in an iterative confirmation bias. An audit of the IEDB, the biggest database in the field, reveals that as of January 2025, 55.8% of assessable data are labeled by computational models rather than verified experimentally. This inflates in silico benchmarks while degrading real-world applicability on new data, effectively making it impossible to objectively test model performance, which can lead to choosing suboptimal solutions and decreasing the chance of any therapys clinical success. In silico simulation shows that iterative data corruption maintains high AUROC while top-of-list retrieval collapses. We reframe epitope discovery as a protein-centric learning-to-rank task and introduce deepMHCflare, a model evaluated exclusively on clean data. deepMHCflare achieves 0.80 Precision@4 on mono-allelic benchmarks versus 0.55-0.65 for gold-standard prediction models. A preclinical cancer vaccine study validated that 2 of the 4 deepMHCflare-nominated peptides were immunogenic, with a third independently confirmed in the literature.

bioinformatics↗