Temporal Data Leakage Inflates Machine Learning Performance in Compostable PBAT/PLA Polymer Biodegradation Modeling
Abstract
This study quantifies how evaluation protocol choice, not model architecture, governs apparent machine learning (ML) performance in poly(butylene adipate-co-terephthalate)/polylactic acid (PBAT/PLA) composite biodegradation modeling. Seven regression architectures spanning linear, regularized, kernel-based, ensemble, and Bayesian nonparametric methods were assessed under four protocols: random splitting, temporal holdout, leave-one-condition-out cross-validation, and temporal shuffle. Under random 80/20 splitting, all seven methods converge to R2 ≥ 0.95; a day-permutation diagnostic reveals that up to ∼96% of this apparent performance (ΔR2 = –0.959) reflects temporal adjacency exploitation, not learned degradation kinetics; this interpretation is confirmed by the immunity of a mechanistic exponential decay reference model (R2 = 0.985, unchanged after permutation). Under temporal holdout, a zero-parameter persistence predictor (R2 = +0.95) outperforms all ML methods (best R2 = +0.16); counterfactual validation across four test windows establishes this advantage as phase-conditional. Variance inflation factor (VIF) analysis reveals that environmental feature contributions cannot be independently estimated under the three-condition experimental design (VIF = ∞; rank-deficient matrix). A six-item diagnostic toolkit applicable to existing datasets without additional experiments is proposed. Evaluation protocol reform, not algorithmic innovation, constitutes the more urgent priority for reliable polymer biodegradation prediction.