bioRxiv · 10.64898/2026.04.24.720637
Housekeeping Gene Expression Normalization in Transcriptomics Mitigates Data Leakage in Machine Learning Models
Abstract
BackgroundInappropriate normalization can lead to data leakage and overfitting in machine learning models. Accurately identifying housekeeping genes (HKGs) can overcome this problem and is crucial for normalizing gene expression data, particularly in RNA-Seq experiments. ResultsFirst, we demonstrate that the gene expression of commonly used HKGs significantly changes over time due to immunosuppressive treatments in transplant recipients. Using large public transcriptomic datasets of kidney transplantation, we developed a pipeline based on the genes coefficient of variation, stability, and Gini coefficient, and identified nine stable and better-suitable HKG candidates. Our results demonstrate that these HKGs improve the robustness and generalizability of machine learning models by minimizing data leakage, as evidenced by superior performance compared to benchmark methods like median ratio normalization and trimmed mean of M values. ConclusionsThis approach enables more accurate comparison of gene expression datasets across different clinical scenarios, improving the reliability of biomarker identification and enhancing personalized treatment strategies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ribas, G. T., Riella, C. V., Guizelini, D., Menegatti Rigo, M., Riella, L. V., Borges, T. J.. 2026-04-24. Housekeeping Gene Expression Normalization in Transcriptomics Mitigates Data Leakage in Machine Learning Models. https://doi.org/10.64898/2026.04.24.720637
Cite the original work for its findings. Save a collection to share your selection of sources.