Preprint
Aug 2026
TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
TrustDABench is introduced, a benchmark that operationalizes two diagnostic questions of LLM reliability and robustness and suggests that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.
Boshen Shi, Yize Liu, Chen Zhao et al.
· 0 citations