The Routing Plateau: Understanding the Accuracy Limits of LLM Routers
This analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals, and collectively fail on queries that require instance-specific routing decisions.
Yi-Fan Lu, Qi-Yue Zhang, Shen-Run Zhang et al.
· 0 citations