JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
Whether a judge returns the same verdict when the same request is worded differently remains largely unexamined, so decoding noise is not charged to wording, and the release lets a reader ask the same of any judge not in the roster.