Resource-Efficient Semantic Retrieval Optimization for Retrieval-Augmented Generation Using FAISS, Milvus, HNSW, and Product Quantization
Abstract
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge, but its performance depends heavily on efficient semantic vector search. This paper presents a deployment-oriented empirical study that systematically compares established backends and ANN configurations under a single, controlled RAG pipeline. Specifically, it evaluates FAISS and Milvus using Hierarchical Navigable Small World (HNSW) graphs and Product Quantization (PQ), combined with optimized chunking, embedding generation, Principal Component Analysis (PCA) for dimensionality reduction, retrieval, re-ranking, and LLM-based generation. Experiments on Natural Questions and TriviaQA show that optimized indexing reduces retrieval latency by 52.4%-79.6%, increases throughput by up to 230%, and reduces memory usage by up to 40%, while maintaining competitive answer quality (F1: 0.88-0.99). These results demonstrate that vector indexing configuration is a critical design factor for efficient and practical RAG deployment.