Widespread use of invalid statistical tests in biomedical machine learning
Cross-validation is routinely used to compare performance in biomedical artificial intelligence. Standard tests ignore correlation across cross-validation folds, inflating false-positive rates. In a PRISMA-guided meta-analysis of 184 studies (impact factor [≥] 15) across 30 biomedical fields, 97% use invalid tests. Among studies with abstract-level claims supported by invalid tests, 59% rely on a spurious comparison: one that loses significance after correcting for this correlation. On average, regaining significance requires a 43% larger effect. Extending these findings to the broader literature, we estimate that spurious comparisons occur in nearly two in three studies and support abstract-level claims in nearly one in three studies. Simulations confirm false-positive rates paradoxically approach 100% when cross-validation is repeated to improve stability. We introduce SHARP, a redesign of cross-validation, which best balances false-positive control and power among 13 benchmarked tests. These results reveal widespread fragility in biomedical artificial intelligence and offer a practical route to valid comparisons.