PolyRAG: A Multi-Agent Multimodal Retrieval-Augmented Generation System for Heterogeneous Document Intelligence
Abstract
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by reducing hallucinations and improving factual accuracy. Most RAG implementations, however, are restricted to a single data modality. Real-world document collections span PDFs, spreadsheets, audio, video and images simultaneously. This paper introduces PolyRAG, where a central routing agent coordinates a multi-agent multimodal RAG framework where specialized agents handle each modality. Semantic embeddings are generated locally via all-MiniLM-L6-v2, and stored in a persistent ChromaDB instance. A three-tier LLM fallback chain - Groq (Llama 3.1 70B), Google Gemini 1.5 Flash, and a local Ollama instance - keeps the user use the system even without internet access. PolyRAG was evaluated on a manually constructed heterogeneous test corpus, achieving a context precision of 1.000, answer similarity of $\mathbf{0. 8 8 3}$, with an end-to-end retrieval latency of $\mathbf{2 3. 4} \text{ms}$.