Selection-aware metrics provide complementary decision-level information in plant genomic prediction
Abstract
Breeding programs increasingly need prediction systems that support reliable candidate selection across heterogeneous environments, which makes cross-environment genomic prediction an important component of multienvironment crop improvement. However, model evaluation still relies heavily on Pearson’s correlation coefficient ( r ), a global summary that may not fully capture which top-ranked candidates would be advanced. Using maize and switchgrass datasets as focused case studies, we compared global predictive ability with phenotype-referenced candidate diagnostics across within-environment and directed cross-environment evaluations of rrBLUP/ridge, Light Gradient Boosting Machine (LightGBM), a multilayer perceptron, and mixed-effect deep neural network (MeNet). Rather than proposing a new selection metric, we used a conditional comparison framework that distinguishes nominal metric reversal from strict decision-level discordance among models with similar Pearson’s r . The r -best and candidate-metric-best models disagreed in 25%–55% of evaluation contexts, but large differences in candidate metrics among near- r models (|Δ r | ≤ 0.02) were uncommon. Nevertheless, models with nearly identical r did not always prioritize the same candidates: LightGBM and rrBLUP/ridge shared, on average, 72.9% of the top 5% candidates in maize and 68.8% in switchgrass. These results indicate that Pearson’s r remains informative for broad model comparison but does not fully describe decision-level behavior. Phenotype-referenced candidate diagnostics are therefore best used as targeted complements when models have similar predictive ability, selection intensity is high, or candidate identity is consequential across environments.