Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matc...