Semi-structured documents are ubiquitous in scientific reports, financial statements, and technical manuals. Question answering over such documents requires simultaneous understanding of text, tables, charts, and complex hierarchical layouts. Existing methods either rely on repeatedly calling large language models for...
TradeLens is introduced, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations, which reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-...
Qiqi Duan, Changlun Li, Chen Wang et al.· arXiv.org· 0 citations
UI2App is the first benchmark targeting interaction inference, the ability to recover application behavior from screenshots alone, without any textual or behavioral guidance, and designs an end-to-end pipeline that evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, a...
Grace Man Chen, Litao Guo, Yifan Wu et al.· arXiv.org· 0 citations
Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.
DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces, is introduced and key challenges for improving data-agent reliability are identified.
Boyan Li, Zhuo-Wen Liang, Yu-Peng Xie et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.