Multimodal models increasingly reach for tools when solving visual tasks (crop, zoom, rotate, brighten), a paradigm known as thinking-with-images. The central challenge is one of perception: tools mostly serve to expose visual evidence, reasoning over that evidence stays in language, and most targets are ones a human c...
A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout...
Jia-Mu Zhang, Tian-Ze Yang, Yu-Cheng Shi et al.· 0 citations
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four wo...
T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.
Jun-Yao Yang, Yu-Cheng Shi, Zhong-Zhi Li et al.· 1 citation
RadFabric is presented, an agentic AI system that orchestrates fourteen specialized open-source CXR analytics models and two Vision-Language Models through a modular protocol that enables explainable, robust diagnoses across common and rare pathologies while facilitating extensibility through additional agents.
Wenting Chen, Yi Dong, Zhaojun Ding et al.· npj Digital Medicine· 3 citations
Recursive Synthetic Terminal Tasks (RST) is presented, a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale and shows no ceiling, indicating that the process can continue well beyond the scale reported here.
Zhongzhi Li, Yucheng Shi, Zongxia Li et al.· 5 citations
This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on...
Zongxia Li, Zhongzhi Li, Yucheng Shi et al.· arXiv.org· 11 citations· ⚡2
By isolating task-specific patterns into independent modules, CRAM mitigates catastrophic forgetting across tasks and boost parameter efficiency, and utilizes adaptive-rank instantiation to identify the capability gap between existing expert capability and new task demands, and dynamically allocate only the necessary p...
Jun Tang, Zhen Xie, Yucheng Shi et al.· arXiv.org· 0 citations
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a d...
Ruhan Wang, Yucheng Shi, Zongxia Li et al.· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.