Preprint
Jul 2026
Reinforcement Learning for Code Optimization
This work makes execution time learnable through three stages: how code is tested, by building DMC-Optim with large optimization tests and a calibrated sandbox; how speed is turned into reward, by composing correctness and speed in the RL environment and using an offline simulator to predict the most promising configurations.
Pierre Chambon, Kunhao Zheng, Juliette Decugis et al.
· 0 citations