Skip to content

UNBIND: UNlearning By INference-time Directional Steering for Code LLMs

Sep 2026 · 0 citations · 35 references
Computer Science

Abstract

Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbf{UNBIND}, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3\% to 99.1\% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188--262 to 0--2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43\% to 6.45\%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.

View source

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#artificial intelligence Review Oct 2025

Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging directions from YOLOv5 through YOLO27, and examines detection, segmentation, depth, classification, pose, oriented detection, tracking, export, quantization, and deplo...

Ranjan Sapkota, Manoj Karkee · 112 citations · ⚡10

BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

A novel threat is unveiled in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base, enabling the attacker to steer the response without altering the user input or modifying the RAG weights.

Jiaqi Xue, Meng Zheng, Yebowen Hu et al. · 109 citations · ⚡8

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Grammar-Aligned Decoding

This paper proposes adaptive sampling with approximate expected futures (ASAp), a decoding algorithm that guarantees the output to be grammatical while provably producing outputs that match the conditional probability of the LLM's distribution conditioned on the given grammar constraint.

Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick et al. · 70 citations · ⚡5

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.