Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment
This work proposes a two-stage generator-in-the-loop alignment framework that consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignm...