Skip to content

Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

Sep 2026 · 1 citation · 30 references
Computer Science

TL;DR

Signed Rescue Routing is introduced, a budgeted routing method that predicts these two events separately and ranks requests by their difference and shows that predicting incremental value, rather than model uncertainty, is a simple and effective objective for efficient LLM cascades.

Abstract

Large language model (LLM) cascades answer easy requests with a small model and escalate selected requests to a larger model. Most routers prioritize examples on which the small model appears uncertain or likely to be wrong. This proxy ignores a decisive fact: escalation is useful only when the large model corrects the small model, and it is harmful when the large model replaces a correct answer with an incorrect one. We introduce Signed Rescue Routing (SRR), a budgeted routing method that predicts these two events separately and ranks requests by their difference. We show that this signed conditional gain is the Bayes-optimal routing score under a fixed escalation budget. SRR requires only the small model's output statistics at deployment and adds a lightweight two-head router. We evaluate SRR with Qwen3-4B and Qwen3-8B on TBD examples from MMLU, HellaSwag, and ARC-Challenge. Across the accuracy-compute curve, SRR reaches an area of TBD, compared with TBD for a learned small-model error predictor and TBD for entropy routing. These results show that predicting incremental value, rather than model uncertainty, is a simple and effective objective for efficient LLM cascades.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference

No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query. We show that collaboration is non-...

Norah Alballa, Wen-Xuan Zhang, Salma Kharrat et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Dynamic LLM Routers are Often Misguided

Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14 settings on a diverse benchmark spanning eight task categories, finding that none of them outperforms a router that randomly selects between two well-c...

Sam Wang, Julia White, Sahibzada Allahyar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing...

Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis · 0 citations
Preprint Aug 2026

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

RLCascadeRouter is a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop''and model selection, and uses trajectory returns and advantages to directly optimize the performance-cost objective.

Shihong Huang, Sheng-Jie Wang, Hong-Yao Ma et al. · 1 citation
#artificial intelligence Preprint Sep 2026

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address...

Wang Wei, Harry Yang, Tiankai Yang et al. · 1 citation
Open access 2026

DABO: Difficulty-Aware Binary Offloading for Collaborative Large-Small Model Inference

DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.

Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.