This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.
Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al.· 0 citations
AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories, is introduced, showing that LLM agents engage in unsafe behavior more often when no safe fulfillment path exists.
Ruoxi Shang, Christina-Maria Androna, Orfeas Menis-Mastromichalakis et al.· 0 citations
The Last Translation Benchmark is introduced, a collection of human-authored and peer-reviewed examples that break leading machine translation models and a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and a...
Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.