bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.08.29.673016

Revealing the Paper Mill Iceberg: AI-Based Screening of Cancer Research Publications.

Abstract

ObjectivesTo train and validate a machine learning model to distinguish paper mill publications from genuine cancer research articles, and to screen the cancer research literature to assess the prevalence of papers that have textual similarities to paper mill papers. DesignMethodological study applying a BERT-based text classification model to article titles and abstracts. SettingRetracted paper mill publications listed in the Retraction Watch database were used for model training. The cancer research corpus was screened by the model, using the PubMed database restricted to original cancer research articles published between 1999 and 2024. ParticipantsThe model was trained on 2,202 retracted paper mill papers and validated on independent data collected by image integrity experts. A total of 2.6 million cancer research papers were screened. Main outcome measuresClassification performance of the model. Prevalence of papers flagged as similar to retracted paper mill publications with 95% confidence intervals and their distribution over time, by country, publisher, cancer type, research area, and within high-impact journals (Decile 1). ResultsThe model achieved an accuracy of 0.91. When applied to the cancer research literature, it flagged 9.87% (95% CI 9.83 to 9.90) of papers and revealed a large increase in flagged papers from 1999 to 2024, both across the entire corpus and in the top 10% of journals by impact factor. Over 170,000 papers affiliated with Chinese institutions were flagged, accounting for 35% of Chinese cancer research articles. Most publishers had published substantial numbers of flagged papers. Flagged papers were overrepresented in fundamental research and in gastric, bone, and liver cancer. ConclusionsPaper mills are a large and growing problem in the cancer literature and are not restricted to low impact journals. Collective awareness and action will be crucial to address the problem of paper mill publications.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Scancar, B., Byrne, J. A., Causeur, D., Barnett, A. G.. 2025-09-03. Revealing the Paper Mill Iceberg: AI-Based Screening of Cancer Research Publications.. https://doi.org/10.1101/2025.08.29.673016

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

BAP1 loss and PRAME expression converge to remodel the tumor-immune ecosystem during uveal melanoma progression

Uveal melanoma (UM) is characterized by a small number of recurrent genetic alterations that determine metastatic propensity. BAP1 loss and PRAME expression define the dominant prognostic axes in UM, yet how they promote malignant progression remains unclear. We profiled 190,535 cells from normal uvea, uveal nevus, primary and metastatic UM using single-cell transcriptomics, T cell receptor sequencing, spatial transcriptomics and isogenic perturbation models. Normal melanocytes, nevus cells and UM cells formed a transcriptional continuum marked by loss of differentiation and emergence of neural crest-like, stress-responsive, hypoxic-glycolytic and immune-interacting states. BAP1 loss and PRAME expression imposed distinct but convergent immunoregulatory programs, inducing interferon and TNF-NFkB signaling and MHC-I expression, with HLA-E showing the strongest response. These alterations were accompanied by macrophage and CD8+ T cell remodeling. PRAME-enriched tumor regions formed spatially organized niches enriched for macrophages and plasma cells. These findings define BAP1 loss and PRAME expression as distinct but convergent axes of tumor-immune coevolution and nominate HLA-E as a candidate mediator of immune resistance.

cancer biology↗

A plasma metabolomics workflow for breast cancer detection using quantitative GC/MS and machine learning

Blood-based metabolomic profiling has been widely investigated for breast cancer (BC) detection; however, clinical implementation remains limited due to variability in sample handling, analytical reproducibility, and overfitting during statistical analysis. We established a plasma GC/MS metabolomics workflow for discriminating BC from healthy controls (HC) using conventional machine-learning algorithms. Plasma samples (n = 360; BC = 180, HC = 180) were collected prospectively under standardized preanalytical conditions before surgery and the initiation of systematic anticancer therapy and analyzed using a quantitative GC/MS platform with automated derivatization. Feature selection and model development were conducted using three machine-learning (ML) algorithms (Lasso logistic regression (LR), random forest classifier (RFC), and support vector machine (SVM)). A total of 45 metabolite candidate biomarkers were identified, and the optimal number of metabolite features for each algorithm was estimated by a recursive feature elimination (RFE)-based strategy. The best-performing models achieved area under the ROC curve values (AUC) of 0.910 (LR), 0.893 (RFC), and 0.843 (SVM). We selected prioritizing candidate biomarkers consistently expressed across the multi-algorithm pipeline. A bagging ensemble model improved stability (AUC = 0.911) and reduced false-positive predictions in the independent HC dataset. In addition, model stability with respect to false-positive predictions was assessed using an independent HC cohort (n = 15) that was collected at a separate institution. These results indicate that a plasma metabolomics workflow combined with conventional multi-algorithm ML, algorithm-specific feature selection, and independent assessment provides stable discrimination between BC and HC in a moderately sized cohort.

cancer biology↗

Prognostic value, signal interaction network, and immune infiltration characteristics of MET gene expression in gastric cancer analyzed by multi-database bioinformatics

Objective Based on the bioinformatics method of multi-database integration, this study systematically analyzes the expression characteristics, clinical pathological correlation, prognostic value, potential molecular mechanisms, and immune infiltration patterns of hepatocyte growth factor receptor (MET) in gastric cancer. Methods The UALCAN and GEPIA databases were employed to examine the differential expression of MET between gastric cancer and normal gastric mucosal tissues, as well as its associations with clinicopathological features. Kaplan-Meier Plotter was utilized to evaluate the impact of MET expression on overall survival (OS) and progression-free survival (PFS). Protein-protein interaction (PPI) network was constructed via LinkedOmics, followed by Gene Ontology (GO) functional annotation and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis of co-expressed genes. Four algorithms, TIMER, CIBERSORT, EPIC, and MCPcounter, were used to cross evaluate the correlation between MET expression and immune cell infiltration; Further validate the cell type specific expression of MET using the gastric cancer single-cell sequencing queues (GSE134520, GSE167297) built into the TISCH database. Results MET expression was significantly elevated in gastric cancer tissues compared with normal gastric mucosa (P < .05), and its expression level was significantly correlated with tumor grade and TNM stage. Patients with high MET expression exhibited significantly poorer OS and PFS than those with low MET expression (P < .05). The PPI network revealed that MET could interact with 20 key proteins, including EGFR, ERBB2, HGF, STAT3, and GRB2 etc. GO enrichment analysis suggests that differentially expressed genes are significantly enriched in functions such as the ERBB signaling pathway, cadherin binding, and DNA repair complexes; KEGG enrichment analysis showed that MET related genes were significantly enriched in pathways such as homologous recombination, nuclear cytoplasmic transport, mismatch repair, and oxidative phosphorylation. Immune infiltration analysis showed that the negative association between MET and B cells infiltration has cross algorithm robustness, while the association with neutrophils, CD8+ T cells, CD4+ T cells, and macrophages exhibits algorithmic heterogeneity or insignificance; There is no significant correlation between MET and common immune checkpoint molecules such as PD-1, PD-L1, CTLA4, etc. Single cell validation further confirmed that MET is mainly enriched in malignant epithelial cells and endothelial cells, and is almost not expressed in immune cells. Conclusions Multidimensional bioinformatic analyses demonstrate that elevated MET expression serves as an independent risk factor for unfavorable prognosis in gastric cancer. MET may mediate dual drug resistance in gastric cancer via crosstalk with multiple signaling molecules (including EGFR, ERBB2, HGF, STAT3 and GRB2) and dysregulation of the homologous recombination repair pathway. Results from multiple-algorithm immune infiltration analysis, single-cell dataset analysis and immune checkpoint correlation analysis indicate that MET exerts only modest direct regulatory effects on the gastric cancer immune microenvironment. This exploratory study offers systematic bioinformatic evidence supporting MET as a candidate prognostic biomarker and potential therapeutic target for gastric cancer. Further functional experiments and prospective cohort studies are required to validate its molecular mechanisms and clinical utility.

cancer biology↗