A novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution, and proposes 11 mutation operators to simulate realistic tool evolution within 123 MCP servers, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
Huanxi Liu, Kun Hu, Jiaqi Liao et al.· 0 citations
This paper examines audio self-supervised learning through the alignment between pretraining objectives, architectural inductive biases, and downstream applications, and relates these demands to the biases of CNNs, recurrent and State Space Models, Transformers, and hybrid architectures.
ReFrame is a training-free multimodal input reframing framework where two agents share a lightweight locally deployed MLLM: the evidence-generation agent constructs complementary risk and utility evidence, and the rewrite-and-routing agent converts it into a safe proxy prompt and image-routing decision before calling the downstream MLLM, without modifying it or accessing its internal information.
Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai et al.· 0 citations