OmegaUse-OfficeVal is introduced, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding, and code-based verifiers from fine-grained rubrics are developed to support stable evaluation.
Abstract
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.
This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings ac...
Kai-Yuan Liu, Qiu-Yang Mang, Bo-Fei Peng et al.· 2 citations
Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113)...
A unified framework that evaluates the capability of models to automate and augment another agent's performance, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Pattaraphon Kenny Wongchamcharoen, K. Gulati, Min-Min Fong et al.· 2 citations
The results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics, and for teams responsible for agent orchestration and reliability at scale to consider horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics.
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs...
Tianyou Wang, Chong-Yang Gao, Ke-Zhen Chen et al.· 1 citation
This work introduces early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task within each task, and instantiates EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features and hal...