Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation....
Wen-Bin Hu, Hui-Hao Jing, Hao-Chen Shi et al.· 0 citations
Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. Although recent studies have developed frameworks to estimate PT parameters for Large Language Models (LLMs...
Rui Wang, Qi-Han Lin, Jiayu Liu et al.· 2 citations
This work proposes soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.
Hui-Hao Jing, Wen-Bin Hu, Shao-Jin Chen et al.· 0 citations
Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V, and proposes PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition.
Baixuan Xu, Yinyui Xu, Tianshi ZHENG et al.· 0 citations
Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
This survey treats isolation as a first-class principle for LLM-agent system safety, and organizes the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment.
This work proposes InferenceDynamics, a flexible and scalable multi-dimensional routing framework by modeling the capability and knowledge of models, and demonstrates its effectiveness and generalizability in group-level routing using modern benchmarks including MMLU-Pro, GPQA, BigGen-Bench, and LiveBench.
Haochen Shi, Tianshi ZHENG, Weiqi Wang et al.· Annual Meeting of the Associ...· 0 citations
This work proposes RLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward, and suggests that code agents can be trained not only to pass tests, but also to optimize the programs they write.
Huihao Jing, Hao-Zhe Cui, Wenbin Hu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.