While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact. To achieve true fo...
Jia-Lu Wang, Pei-Zhi Niu, Hao-Teng Yin et al.· 0 citations
Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and de...
Xin-Jie Shen, Jun-Ran Wang, Rong-Zhe Wei et al.· 0 citations
The Rashomon Explanation paradigm is introduced, which builds a set of faithful, prediction-guiding explanations rather than a single one, and it is proved that this set is generally non-empty and that explanation fidelity bounds the performance of the models it guides.
Pan Li· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.