Skip to content

ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation

Sep 2026 · 1 citation · 28 references
Computer Science

TL;DR

Overall, dense retrieval gives the best accuracy on this dataset, and additional reranking adds processing time without improving classification performance, while Category-level analysis shows strong performance on Shipping, Cancellation, and Return, while Unknown queries remain the main source of errors.

Abstract

Enterprise customer support systems must answer customer questions correctly, retrieve the right policy information, use customer context, and pass difficult cases to human agents when needed. This paper presents ShopEase, a Generative AI-based multi-agent framework for enterprise customer support. The system combines six components: Intent, CRM, Memory, Hybrid RAG, Escalation, and Supervisor, and uses LLaMA 3.2 running locally through Ollama for response generation. The retrieval module combines FAISS (dense retrieval) and BM25 (sparse retrieval), and six configurations are evaluated: BM25-only, FAISS-only, Fair RRF, Weighted RRF, RRF with Cross-Encoder, and Top-10 Hybrid with Cross-Encoder. Instead of using a fixed mapping between intent and policy, the policy category is decided directly from the retrieved documents. The system was evaluated on 2632 held-out customer queries across six categories: Refund, Return, Shipping, Cancellation, Damaged Product, and Unknown. FAISS-only achieved the highest accuracy of 85.37\% (2247 correct predictions), closely followed by Weighted RRF at 85.07\%. BM25-only achieved only 55.74\% accuracy. Adding cross-encoder reranking did not improve results: RRF with Cross-Encoder reached 83.24\%, and Top-10 Hybrid with Cross-Encoder reached 81.88\%, while also increasing response latency. Category-level analysis shows strong performance on Shipping, Cancellation, and Return, while Unknown queries remain the main source of errors. Statistical testing using McNemar's test shows no significant difference between FAISS-only and Weighted RRF, though both perform significantly better than Fair RRF and the cross-encoder configurations. Overall, dense retrieval gives the best accuracy on this dataset, and additional reranking adds processing time without improving classification performance.

View source

Similar papers

Conference Aug 2026

Hybrid Retrieval-Augmented Generation and Multi-Agent AI Systems for Explainable Enterprise Automation

With the increasing use of AI in the business world, there is a strong demand for reliable, knowledgeable and explainable automated systems. Large Language Models (LLMs) could perform language-based processes automatically, but the risks of hallucination, stale knowledge and opaqueness make them less practical in a ser...

Praveen Dommalapati · 0 citations

BizSage: A Self-Evolving Multi-Agent Framework for Business Research with Efficient Knowledge Retrieval

BizSage is presented, a multi-agent framework combining corpus-level fine-grained retrieval with quality-driven self-evolution that paves the way for reliable research assistance in economics, business, and the broader social sciences.

Yu-He Wu, Guang-Yu Wang, Jia-Xin Liu et al. · 0 citations
Open access Sep 2026

A Hybrid BDI + RAG Multi-Agent Architecture: Plan-Based Reasoning over Multimodal Retrieval

Retrieval-Augmented Generation (RAG) gives language models access to external text and image collections, but it does not make the decisions taken on that content reproducible. Classical Belief–Desire–Intention (BDI) agents provide explicit plans and traceable decisions, although they normally expect beliefs in symboli...

Halil Yesil, Baris Tekin Tezel, Moharram Challenger · 0 citations
Open access 2026

Creating an AI Agent Using Unique Language Attributes

This paper presents a Hebrew-first local LLM chat agent that combines Retrieval-Augmented Generation (RAG), citation-aware document answering, controlled web search, and full right-toleft (RTL) user interaction. Unlike cloud-only assistants, the default response path operates locally, supporting privacy, predictable op...

Michael Sirkovich, M. Domb · 0 citations
Conference Aug 2026

Designing Persistent Intelligence in Marketing Agents: A Hybrid Memory–Retrieval Framework for Context-Aware Knowledge Generation

Large language models that enable enterprise marketing agents must have both correct knowledge underpinning and ongoing contextualization to provide quality and personalized experiences. Traditional Retrieval-Augmented Generation (RAG)is an effective system that enhances the factual accuracy, by using external sources...

Raunak Kumar, Ankur Gupta, Naveen Kumar Mylarappa et al. · 0 citations
#graph neural networks Open access Aug 2026

Adaptive Decision Intelligence through Retrieval-Augmented Multi-Agent Large Language Models

A comprehensive five-layer framework comprising foundation benchmarks, dynamic hybrid retrieval, multi-agent collaboration with weighted consensus, knowledge graph evolution through graph neural networks, and adaptive human-AI interaction is proposed, establishing a robust foundation for trustworthy, scalable, and trul...

Manish Rana · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.