Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they...
Chang-Liang Zhou, Yuan-Yao Chen, Rong-Sheng Chen et al.· 0 citations
GraphSkillEvo is introduced, a population-based evolutionary optimization framework with mutation and crossover operators for graph-structured skills that enables broader and more comprehensive exploration of the structured skill space than purely LLM-based iterative self-refinement.
Rui Sun, Zhi Zheng, Zhen-Kun Wang et al.· 0 citations
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs...
Zhi Zheng, Rong-Sheng Chen, Yunpeng Ba et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.