The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints
Standard accuracy flatters all ten models the authors evaluate and reorders their ranking, and prompting fixes that look effective largely vanish under paired scoring.
2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Standard accuracy flatters all ten models the authors evaluate and reorders their ranking, and prompting fixes that look effective largely vanish under paired scoring.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and re...
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.