Skip to content
Review

Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching

Sep 2026 · 1 citation · 72 references
Computer Science

TL;DR

This work identifies a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill.

Abstract

Skill selection is a key stage in LLM-agent workflows, determining which installed skill should handle a user request. Existing attacks on this stage primarily rely on explicit prompt injection or instruction-level steering, which can expose recognizable manipulation signals. In this work, we identify a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill. Based on this observation, we present Implicit Skill-Selection Manipulation via Semantic Matching (ISM), which jointly shapes target-skill metadata and reusable prompts to manipulate skill selection without explicit selection instructions. Specifically, we develop a three-stage strategy to broaden semantic coverage, strengthen target distinctiveness, and preserve natural prompt wording. Across four task domains and eight selector models, ISM increases the average target-selection rate (TSR) from 15.2% to 63.5%. In a matched comparison, ISM achieves a 73.5% TSR, only 9.8 percentage points below Explicit Steering. Human reviewers block ISM in only 2.9% of judgments, versus 91.4% for Explicit Steering, while five LLM-based inspectors pass ISM at an average rate of 82.9%, versus 37.4% for Explicit Steering. Moreover, ISM remains effective against PPL-W, Llama Prompt Guard 2, and PIGuard.

View source

Similar papers

Preprint Aug 2026

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

SkillReason-Bench is introduced, a large-scale cross-domain benchmark containing 3,729 queries and a retrieval corpus of 61,228 skills spanning nine domains and SkillRea- son is proposed, a two-stage framework that uses chain-of-thought rea- soning as training-time supervision for skill retrieval.

Donghong Jiang, Endian Lin, Luoping Cui et al. · 2 citations
Preprint Sep 2026

A Finger on the Scale: Covert Policy Steering through Agentic Skills

SkillShift is presented, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking that combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validi...

Jia-Rui Li, Jia-Hao Chen, Chun-Yi Zhou et al. · 0 citations
Preprint Aug 2026

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Skill-Use is introduced, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it.

Jinyi Han, Yuanjian Xu, Ying Liao et al. · 4 citations · ⚡2
#artificial intelligence Preprint Sep 2026

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated re...

Guanqun Yang, Wen-Long Zhang, Tian Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillAlign: Aligning Skill Interfaces for LLM-based Agents

SkillAlign is proposed, a provider-agnostic framework that represents candidate skills as multi-view procedural cards and renders them through alternative exposure interfaces, including full instructions, hints, compressed summaries, workflows, or no exposure, which enables counterfactual evaluation where the task, age...

Shuo Ren, Xiaomian Kang, Jia-Jun Zhang · 0 citations
Preprint Aug 2026

SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion

SkillCommit is an online skill evolution framework that continuously transforms experience into a hierarchical library of reusable skills, enabling cross-model experience transfer and consistently improves agent performance across diverse domains.

Yu He, Wei-Kai Yang · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.