K-Bench is introduced, a benchmark that scores LLM unlearning under agentic deployment and certifies forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten.
Guang-Sheng Yu, Yan-Na Jiang, Qin Wang et al.· 0 citations
Traffic Signal Control (TSC) is a safety-critical cyber-physical system that relies on real-time sensing. Corrupted observations caused by adversarial perturbations or sensor failures can propagate from the sensing layer into the controller and degrade traffic efficiency. Existing robust Reinforcement Learning (RL)-bas...
Ming-Yuan Li, Chun-Yu Liu, Xiao Liu et al.· 0 citations
Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An of...
Ming-Yuan Li, Yan-Na Jiang, Guang-Sheng Yu et al.· 0 citations
Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent. We introduce K-Bench, a benchmark that scores L...
Guangsheng Yu, Yanna Jiang, Qin Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.