While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systemat...
Zong-Wan Cao, Zi-Yuan Yang, Shang-Bin Feng et al.· 0 citations
A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on whic...
Christina Hahn, Shang-Bin Feng, Dean Light et al.· 0 citations
Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through...
Ji-Han Yao, Si-Han Zeng, Shang-Bin Feng et al.· 0 citations
FLIP (FLipped Inference for Prompt reconstruction), a reference-free and rubric-free reward modeling approach that reformulates reward modeling through backward inference that enables reliable reward modeling in downscaled regimes where judgment methods fail, is proposed.