FinancialAuditBench is introduced, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements for model evaluation and training in privacy-sensitive domains.
Jerry Huang, Sarvesh Babu, Matt Van Buren et al.· 0 citations
V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions, is introduced, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions.
Si-Qi Zeng, Sewoong Lee, Han Zhao et al.· arXiv.org· 2 citations
ReasoningFlow is introduced, a framework that captures the discourse structures of LRM reasoning traces into fine-grained directed acyclic graphs (DAGs) and reveals diverse fine-grained reasoning behaviors that can be used for better reasoning trace monitorability.
The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, providing a safe and controllable experimental environment for assessing LLMs'ability to perform realistic construction engineering tasks.
Sewoong Lee, Risham Sidhu, J. Hockenmaier et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.