Sep 2026· Asia Pacific Economic and Management Review· 0 citations· 46 references
TL;DR
This work formalizes CIGD via the drift amplification ratio (DAR), evaluates six compression strategies on a controlled injection benchmark ( per cell), and proposes Goal-Anchored Compression (GAC): a pinned goal anchor, a negation ledger, status-aware retention scoring, and drift re-anchoring.
Abstract
Long-horizon LLM agents rely on context compression to keep growing histories within a finite window [1][2]. These mechanisms are goal-agnostic—retention follows volume, recurrence, and generic salience rather than task relevance—and we show they systematically amplify off-task content, inducing Compression-Induced Goal Drift (CIGD). We formalize CIGD via the drift amplification ratio (DAR), evaluate six compression strategies on a controlled injection benchmark ( per cell), and propose Goal-Anchored Compression (GAC): a pinned goal anchor, a negation ledger, status-aware retention scoring, and drift re-anchoring. Volume- and recurrence-based compressors amplify detours monotonically with repetition (attention-scored eviction reaches at twelve-fold repetition), while query-conditioned compressors de-amplify. Under severe budgets, summarization and scored eviction lose the user’s negated constraint in and of episodes where GAC retains it in ; GAC eliminates residual detour content () with smaller contexts. Prior work asks what compression forgets [3][4]; we ask what it amplifies.
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training c...
Shantanu Dixit, Anson Bastos, Xu-Chao Zhang et al.· 0 citations
This work introduces two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask, and proposes SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation.
J. Zinco, Xun-Jie Zhu, Shen Huang et al.· 0 citations
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle''phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, methods based on fixed...
Zi-Yi Zhang, Shuang Cui, Hao-Tian Zhang et al.· 0 citations
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference t...
Sonia Laguna, João Monteiro, Marco Cuturi et al.· 1 citation
Decay-Aware State Compression (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout to integrate efficiently with tensor-parallel inference engines.
Yanzhi Yu, Ping-Wei Sun, Jian-Chao Tan et al.· 2 citations
Results motivate CertKV, a training-free compressor that reserves one tail-summary slot per head and allocates the rest by value dispersion, which is top-two in seven of nine LongBench-v2 settings, remains in the leading compressed tier on 128K RULER, and realizes a ten-fold cache budget in a packed Llama prototype.
Chi-Wun Yang, Xiaoyu Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.