bioRxiv Science⌕ Search

Biology subjects

Koh, J. M. S.

Publications and source records attributed to Koh, J. M. S..

2 recordsLinked to original sources

Federated deep learning enables cancer subtyping by proteomics

Artificial intelligence applications in biomedicine face major challenges from data privacy requirements. To address this issue for clinically annotated tissue proteomic data, we developed a Federated Deep Learning (FDL) approach (ProCanFDL), training local models on simulated sites containing data from a pan-cancer cohort (n=1,260) and 29 cohorts held behind private firewalls (n=6,265), representing 19,930 replicate data-independent acquisition mass spectrometry (DIA-MS) runs. Local parameter updates were aggregated to build the global model, achieving a 43% performance gain on the hold-out test set (n=625) in 14 cancer subtyping tasks compared to local models, and matching centralized model performance. The approachs generalizability was demonstrated by retraining the global model with data from two external DIA-MS cohorts (n=55) and eight acquired by tandem mass tag (TMT) proteomics (n=832). ProCanFDL presents a solution for internationally collaborative machine learning initiatives using proteomic data, e.g., for discovering predictive biomarkers or treatment targets, while maintaining data privacy. Statement of SignificanceA federated deep learning approach applied to human proteomic data, acquired using two distinct proteomic technologies from 40 tumor cohorts from eight countries, enabled accurate cancer histopathological subtyping while preserving data privacy. This approach will enable privacy-compliant development of large-scale proteomic AI models, including foundation models, across institutions globally.

cancer biology↗

Heat n Beat: A universal high-throughput end-to-end proteomics sample processing platform in under an hour

Proteomic analysis by mass spectrometry (MS) of small ([≤]2 mg) solid tissue samples from diverse formats requires high throughput and comprehensive proteome coverage. We developed a near universal, rapid and robust protocol for sample preparation, suitable for high-throughput projects that encompass most cell or tissue types. This end-to-end workflow extends from original sample to loading the mass spectrometer and is centred on a one tube homogenisation and digestion method called Heat n Beat (HnB). It is applicable to most tissues, regardless of how they were fixed or embedded. Sample preparation was divided to separate challenges. The initial sample washing, and final peptide clean-up steps were adapted to three tissue sources: fresh frozen (FF), optimal cutting temperature (OCT) compound embedded (FF-OCT), and formalin-fixed paraffin-embedded (FFPE). Thirdly, for core processing, tissue disruption and lysis were decreased to a 7 min heat and homogenisation treatment, and reduction, alkylation and proteolysis were optimised into a single step. The refinements produced near doubled peptide yield, delivered consistently high digestion efficiency of 85-90%, and required only 38 minutes for core processing in a single tube, with total processing time being 53-63 minutes. The robustness of HnB was demonstrated on six organ types, a cell line and a cancer biopsy. Its suitability for high throughput applications was demonstrated on a set of 1,171 FF-OCT human cancer biopsies, which were processed for end-to-end completion in 92 hours, producing highly consistent peptide yield and quality for over 3,513 MS runs. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=152 SRC="FIGDIR/small/559846v2_ufig1.gif" ALT="Figure 1"> View larger version (34K): org.highwire.dtl.DTLVardef@4b8399org.highwire.dtl.DTLVardef@1acc573org.highwire.dtl.DTLVardef@1d7155eorg.highwire.dtl.DTLVardef@1bbef10_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗