BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.