Preprint
Aug 2026
SEER: Long-Context Reasoning via Selective Visual-Text Compression
SEER is presented, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning.
Jiawei Xu, Zhilin Zhai, Jinrui Fang et al.
· 0 citations