As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based met...
Xu-Kai Wang, Liangqi Li, Zhi-Yu Xu et al.· 0 citations
This work introduces PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment, and proposes Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision al...
Jun-Chi Chen, Chang-Tao Miao, Yu Xiang et al.· 0 citations
DT-Guard is presented, a content safety guardrail model based on a Reasoning-Active Training, Reasoning-Free Inference paradigm, which demonstrates that reasoning supervision can be effectively internalized into low-latency safety discrimination.
Heng-Xiang Liu, Changtao Miao, Xinjie Yang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.