Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves sol...
Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al.· 0 citations
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We inves...
Minki Kang, Ryo Hachiuma, Shao-Kun Zhang et al.· 0 citations
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple...
Chanuk Lee, Minki Kang, Sangwoo Park et al.· 0 citations
Experiments show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings, and suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.
Sehee Kim, Yumin Choi, Minki Kang et al.· 0 citations
Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development, is introduced, which outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.