With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing addresses this need by dynamically assigning queries to the m...
Yao Lu, Zhai-Yuan Ji, Ya-Xin Gao et al.· 0 citations
Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory...
Xuanze Chen, Xukang Xie, Wentao Fu et al.· arXiv.org· 0 citations
This work introduces AgentS4D, a sandboxed benchmark for lifecycle-wide runtime safety evaluation and evaluates all 20 combinations of four harnesses and five LLM backends, finding that the observed safety of an agent system varies with both its harness-LLM pairing and how risk is introduced.
Jiajun Zhou, Zhaoxuan Ke, Jihang Ye et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.