How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.
Yue Huang, Zhangchen Xu, Yu-Chen Ma et al.· 1 citation
Long-term memory enables personalized conversational agents to retain user information across sessions. However, existing memory architectures primarily optimize for utility while neglecting the risks of unnecessarily storing and reusing private attributes such as personally identifiable information (PII). Addressing p...
Wen-Jie Wang, Wen-He Si, Xinyue Xu et al.· 1 citation
Safety Sentry is instantiated, a lightweight guard model whose inference reduces to a single decoding call that outperforms a broad set of open-weight and frontier closed-source baselines on overall accuracy and safety-related recall, while controlling both directional error rates simultaneously.
NapMem is introduced, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context, and suggests that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.
MemoHarness is introduced, an adaptive harness optimization framework that learns from its own executions and improves over the fixed harnesses it is compared against and shows selective transfer to unseen suites and base models.
Yue Huang, Wenjie Wang, Han Bao et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.