A Secure Multi-Tenant Framework for Institutional Question Answering with Tool-Calling LLMs
Abstract
Large language model (LLM)-based question-answering (QA) systems are rapidly being deployed in institutional environments, where queries often require both unstructured document retrieval and structured database access. This paper proposes a secure multi-tenant framework for institutional QA that unifies tool-calling LLMs, document retrieval, SQL/BI querying, and security guardrails within a single pipeline. The proposed framework classifies queries based on user context and selectively or jointly invokes document retrieval and SQL tools, applying security guardrails at the response stage. We evaluate the framework on two benchmarks: an in-house multi-tenant security benchmark based on real institutional data from the Department of Liberal Studies (DLS) at the Catholic University of Korea, and the public Spider Text-to-SQL benchmark for external SQL pipeline validation. On the in-house benchmark, the full system (Configuration C) achieved 89.46% overall accuracy and, on 108 restricted and adversarial queries, complete refusal correctness, zero tenant information leakage, and 100% security accuracy—demonstrating that the security layer functions as an explicit architectural control rather than an emergent property of retrieval accuracy. On the Spider benchmark, the schema-grounded SQL tool configuration achieved 73.31% execution accuracy without benchmark-specific optimization, providing evidence that the SQL pipeline can be executed in an external cross-domain setting. These results demonstrate that SQL tool-calling and security guardrails are both essential for performance and safety in multi-tenant institutional QA.