Skip to content

Author

Zhong Ji

We have 2 of 29 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering

Long-document multimodal question answering requires more than retrieving relevant chunks from a large document. Different queries require different evidence behavior. Existing multimodal RAG systems improve evidence access through text chunks, page images, graph links, or heterogeneous document elements, but they often apply a largely query-agnostic evidence-use strategy. We present TAP-RAG, a task-aware policy-controlled RAG framework for long-document multimodal QA. TAP-RAG contains a main controller, the Task-Aware Policy Controller (TAPC), and two policy-guided evidence executors: Task-Aware Query-Guided Flow Diffusion (TA-QFD) and Task-Aware Visual Enhancement (TAVE). For each query, TAPC predicts the task prior, estimates visual/local/global evidence signals, and produces an executable policy. TA-QFD then expands textual and structural evidence over the multimodal document graph, while TAVE selectively inspects page images when visual or layout evidence is needed. A guarded synthesis stage fuses text, visual, and structural evidence and abstains when support is insufficient. On DocBench and MMLongBench-Doc, TAP-RAG achieves the best overall accuracy among the compared systems, improving over a matched multimodal-RAG baseline by +9.1 points (61.1 to 70.2) and +4.5 points (42.2 to 46.7), respectively.

Zhong Ji, Keqi Jin, Yan Zhang et al. · 0 citations
Preprint Jul 2026

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

CaRe is proposed, a training-free robust framework that calibrates compact visual representations before reasoning to preserve semantic fidelity in VLMs and outperforms state-of-the-art token reduction baselines.

Jiasheng Li, Zhong Ji, Yan Zhang et al. · 0 citations