This analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals, and collectively fail on queries that require instance-specific routing decisions.
Yi-Fan Lu, Qi-Yue Zhang, Shen-Run Zhang et al.· 0 citations
Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture...
Yi-Fan Lu, Qi-Yue Zhang, Haotian Shan et al.· 0 citations
This work introduces a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods and shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robus...
Jia-Mu Zhang, Ling-Xi Zhang, Peng-Jun Lu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.