JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative jud...