Preprint
Jul 2026
SALT: Salience-Aware Lexical Trie for Long-Context Compression
SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure, reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that target decoding-time latency and memory.
Oteo Mamo, Hyunji Yi, Joydhriti Choudhury et al.
· 0 citations