This work proposes a two-stage generator-in-the-loop alignment framework that consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignment.
Abstract
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce a correct answer. Motivated by this, we propose a two-stage generator-in-the-loop alignment framework that closes this gap without human document-level relevance annotations. Our framework consists of two stages: in Stage 1, a VLM generates a hypothetical text passage from the image-query pair, which is used as the retrieval query for dense text search, bridging the image-to-text modality gap. In Stage 2, a cross-encoder reranker adapted with low-rank adaptation (LoRA) is fine-tuned using answer-supervised preference pairs mined from the frozen VLM: given the dataset answer label, a candidate document is labeled positive if the VLM produces the correct answer when given that document as context, and negative otherwise. This generator-guided signal is compatible with multiple alignment loss functions, including contrastive (triplet) loss, pairwise direct preference optimization (DPO), and supervised fine-tuning (SFT), and supports periodic re-mining to refresh preference pairs as the reranker improves. Experiments on VQA-X and A-OKVQA with Qwen3.5-2B and Qwen3-VL-4B-Instruct show that our proposed framework consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignment.
This study proposes STaR, a novel retriever fine-tuning framework that integrates BM25 similarity graph-based soft labeling with a triplet similarity learning strategy based on Sentence-BERT (SBERT), and introduces a triplet-aware SBERT training architecture that explicitly models relative semantic distances between qu...
Jiali Jiang, Chih-Yung Chang, Youxi Li et al.· Multimedia Systems· 0 citations
Retrieval-augmented generation (RAG) has attracted significant attention for enhancing large language models (LLMs) in domain-specific and knowledge-intensive tasks by utilizing external documents retrieved by retrievers. However, LLMs often struggle to determine which retrieved documents are relevant and how they rela...
Fu-Da Ye, Shuang-Yin Li, Yong-Qi Zhang et al.· ACM Transactions on Knowledg...· 0 citations
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
MCoRe, a multi-entry complementary retrieval framework with reflection-guided iteration for multi-hop QA that enables multi-entry complementary retrieval by indexing entry units at multiple semantic resolutions with explicit links to chunk evidence, and fusing cross-resolution hits via chunk-level voting to form a comp...
Ju-Xiang Zeng, Zhuohui Gao, Zhe Hou et al.· Proceedings of the 32nd ACM...· 0 citations
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on h...
Tzu-I Ho, Yung-Yu Shih, Shang-Yu Su et al.· 0 citations
Large language models (LLMs) have substantially advanced natural language generation (NLG), but high-quality content creation remains constrained by hallucination, limited controllability, retrieval noise, domain shift, and computational cost. This study proposes Retrieval-Augmented and Preference-Aligned Generation (R...
Jian Zhang, Yuxuan Zheng, Feng-Jun Ye et al.· Archives des sciences: a mul...· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.