Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al.· 0 citations
SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.
Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al.· 0 citations
This work introduces OR-Clarify, a benchmark for pre-formulation clarification and proposes Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop.
Si-Han Ge, Yi-Chen Lin, Chen-Yu Zhou et al.· 0 citations
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shapes which trajectories are reachable. We study post-violation recovery admission, where p...
Chen-Yu Zhou, Qi-Liang Jiang, Shu-Ning Wu et al.· 1 citation
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction o...
Chen-Yu Zhou, Qi-Liang Jiang, Shu-Ning Wu et al.· 0 citations
It is shown that conditioned on a candidate, a judge scores plausibility, not correctness, leaving false-positive basins a policy learns to exploit, and a falsifiable bound predicts which regimes are exposed.
Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers. Cert...
Chenyu Zhou, Qiliang Jiang, Shuning Wu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.