Doc2DB-Bench is introduced, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells, which provides a testbed for reliable, auditable, and relationally faithful LLM-based data systems.
Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang et al.· 0 citations
DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces, is introduced and key challenges for improving data-agent reliability are identified.
Boyan Li, Zhuowen Liang, Yupeng Xie et al.· 1 citation
Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.
Tian Zeng, Yuntao Hong, Zhongjun Ding et al.· 1 citation