CHQR: A Context-Aware Hybrid Query Routing Framework for Energy-and Latency-Optimized Academic Question Answering
Abstract
The rapid adoption of Large Language Models (LLMs) for academic question answering has led to significant computational overhead, increased latency, and high energy consumption, even for queries that require minimal reasoning. Existing systems typically route all queries to LLMs without considering their complexity or information requirements, resulting in inefficient resource utilization. This study proposes Context-Aware Hybrid Query Routing (CHQR), a multi-layer intelligent framework that dynamically routes academic queries across four processing levels: rule-based retrieval, Small Language Models (SLMs), retrieval-augmented generation (RAG), and LLM-based reasoning. The proposed system incorporates a context-aware query analysis module that evaluates query intent, domain, and complexity, along with a confidence-based escalation mechanism that ensures response reliability while reducing computational costs. A cost-aware routing function is formulated to balance accuracy, latency, and energy consumption, extending prior routing schemes with explicit query-complexity and confidence signals. Experimental results on a mixed academic query dataset show that the proposed CHQR framework reduces response latency by up to 37.5% and estimated energy consumption by approximately 30% relative to an LLM-only baseline while maintaining comparable accuracy. These energy values are based on a proxy metric that approximates energy usage from model invocation frequency and execution time rather than direct hardware-level power measurements. The results highlight the effectiveness of demand-aware routing strategies in building efficient and scalable AI-driven academic assistance systems while demonstrating the potential of context-aware model orchestration as a sustainable approach for large-scale AI deployment.