PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle''phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, methods based on fixed...