Preprint
Aug 2026
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
It is shown that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures, and BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural.
Shailja Thakur, Sungeun An, Chad DeLuca et al.
· 0 citations