Multi-agent reasoning systems in high-stakes domains must be both accurate and safe, yet agents often follow heterogeneous value priorities (e.g., rigor, conciseness, safety), causing conflicting recommendations. Existing methods do not jointly provide: (i) principled inference of each agent's implicit values from beha...
Yi-Yao Zhang, Diksha Goel, Hussain Ahmad et al.· 2 citations
Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts obje...
Caoliwen Wang, Meng-Di Wang, Heng Zhang et al.· 0 citations
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four GRPO policies, an unconstrained scalar DeepSeek-...
Yi-Yao Zhang, Diksha Goel, Hussain Ahmad et al.· 0 citations
CausalNav, a controller built around a signed, action-conditioned transition graph over identified state coordinates, is studied with CausalNav, a controller built around a signed, action-conditioned transition graph over identified state coordinates, which attains the best average rank.
Yi-Yao Zhang, Diksha Goel, Hussain Ahmad et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.