DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
The DRAGON dataset contains 11,664 annotated question instances from six diagram QA datasets, with a 2,445-instance test set carrying human-verified evidence annotations and a standardized evaluation framework, which supports future research on models that ground their predictions in visual evidence.