bioRxiv Science⌕ Search

Biology subjects

Savchenko, M.

Publications and source records attributed to Savchenko, M..

2 recordsLinked to original sources

Benchmarking of bulk transcriptomic harmonization tools in a multi-platform B-cell lymphoma cohort identifies feature-specific quantile normalization and surrogate variable analysis as top-performing methods

Cross-platform harmonization of bulk transcriptomic datasets remains a fundamental challenge for developing cancer biomarkers because of persistent unresolved batch effects. Most harmonization tools are benchmarked on datasets with large inter-group biological differences (for example TCGA tumor types), whereas actionable biomarker mining requires preserving subtle transcriptional distinctions between closely related diagnoses. Here we present ComboBatch, a benchmarking pipeline that evaluates the full cross-product of 14 batch-removal strategies, 3 imputation methods, 33 harmonization algorithms and 2 post-removal conditions across 7,174 samples from 88 germinal-center B-cell lymphoma cohorts spanning four transcriptomic platforms. Scoring 87 quality metrics across 2,234 harmonization approaches, we show that method choice (R2 0.36) and batch-removal strategy (0.26) are the principal determinants of harmonization quality, whereas imputation (0.016) and post-removal (<0.01) are secondary. Feature Specific Quantile Normalization and Surrogate Variable Analysis were the top methods, jointly resolving follicular lymphoma, diffuse large B-cell lymphoma and normal germinal-center B-cell differences in multi-platform and RNA-seq-only compositions, respectively. We provide a data-driven five-scenario decision tree for harmonization method selection, applicable to any retrospective multi-platform transcriptomic study. The ComboBatch pipeline is available on GitHub and can be used for harmonization, allowing bioinformaticians to utilize 33 harmonization and 3 imputation methods according to their needs.

cancer biology↗

AI-based separation of malignant cell- and microenvironment-specific gene expression from bulk RNA sequencing enhances biomarker interpretation

Bulk RNA sequencing (RNA-seq)-based gene expression analysis is a promising tool for personalized cancer diagnostics, disease monitoring, and treatment decision-making. However, its clinical utility is limited by interference from non-malignant tumor microenvironment cells, which can dominate transcript data in low-purity tumors. While cell deconvolution methods like Kassandra can predict digital cell percentages from bulk RNA-seq, approaches for delineating the gene expression contribution of tumor compartments remain limited. To overcome this limitation, we developed Helenus, a machine-learning-based tool that separates gene expression between malignant and non-malignant cells. Trained on over 200 million synthetic RNA profiles representing diverse tumor types and purities, Helenus demonstrated high accuracy in separating gene expression origin. Helenus also uncovered true genomic-RNA correlations such as copy number alterations and the expression of therapeutic antibody-drug conjugate targets specifically on tumor cells. Helenus provides critical insights into tumor biology and immunotherapy response by precisely identifying biomarker expressions, paving the way for more effective personalized cancer care. SignificanceHelenus extracts gene expression profiles of cancerous and non-cancerous compartments of tumor biopsies from bulk RNA-seq data, enabling the determination of how the expression of specific genes affects malignancy and tumor immunity.

cancer biology↗