Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...
EvoOntology is introduced, a self-evolving ontology layer for data agents that encapsulates the ontology as an MCP server comprising a schema layer, a content layer, and a tool layer, enabling agents to actively query and interact with the ontology at runtime.
Mei-Duo Chong, Shao-Lei Zhang, Ju Fan et al.· 0 citations
SkillAdam, an Adam-inspired framework for optimizing discrete and non-differentiable skill documents, is introduced, an Adam-inspired framework for optimizing discrete and non-differentiable skill documents with substantially fewer optimization iterations and lower cost than prior methods.
Gao-Yuan Li, Mei-Hao Fan, Yi-Zhe Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.