Data leakage and measurement error inflate the apparent predictability of overyielding from plant traits
Predicting plant mixture overyielding from functional traits is central to understanding how biodiversity influences ecosystem functioning. Combining empirical data from a large mixture experiment (764 mixtures) and two trait-measurement experiments containing 90 soybean genotypes, with complementary simulations where trait-function relationships and noise levels were defined a priori, we show that apparent model performance depends critically on how predictive models are validated. When training and testing data share genotypes, predictive ability is strongly inflated because shared monoculture and trait measurements create data leakage. Measurement errors further propagate through these shared components, inducing spurious correlations and amplifying noise. When validation uses completely independent genotypes, the predictive power of both linear and machine-learning models declines sharply, revealing limited but genuine predictability. These results show how data structure and measurement error can produce misleading model performance and underscore the need for rigorous validation to achieve robust ecological prediction.