Skip to content

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Jul 2026 · arXiv.org · Vol abs/2607.27155 · 1 citation · 43 references
Computer Science

TL;DR

OmegaUse-OfficeVal is introduced, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding, and code-based verifiers from fine-grained rubrics are developed to support stable evaluation.

Abstract

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.

View source

Similar papers

#natural language process... Preprint Sep 2026

When Agents Slow Down: Understanding LLM Agents'Test-Time Strategies via Elo-per-token Analysis

This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings ac...

Kai-Yuan Liu, Qiu-Yang Mang, Bo-Fei Peng et al. · 2 citations
#artificial intelligence Preprint Sep 2026

DeltaSelect: Affordable A/B Testing for Coding Agents

Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113)...

Nicholas J. Conn · 0 citations
#artificial intelligence Case report Open access Aug 2026

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

A unified framework that evaluates the capability of models to automate and augment another agent's performance, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.

Pattaraphon Kenny Wongchamcharoen, K. Gulati, Min-Min Fong et al. · 2 citations
#artificial intelligence Preprint Aug 2026

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

The results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics, and for teams responsible for agent orchestration and reliability at scale to consider horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics.

Shubhra Mittal · 0 citations
#artificial intelligence Preprint Aug 2026

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs...

Tianyou Wang, Chong-Yang Gao, Ke-Zhen Chen et al. · 1 citation
#natural language process... Preprint Sep 2026

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

This work introduces early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task within each task, and instantiates EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features and hal...

Yu-Ling Shi, Zhensu Sun, Jun-Sen Dong et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.