TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
TrustDABench is introduced, a benchmark that operationalizes two diagnostic questions of LLM reliability and robustness and suggests that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.