Clinical selectivity and failure modes of automated chest radiograph report evaluation metrics: a cross-dataset analysis of ReXErr-v1 and RadEvalX
Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert er...