OmegaUse-OfficeVal is introduced, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding, and code-based verifiers from fine-grained rubrics are developed to support stable evaluation.
Jing-Bo Zhou, Yu-Sai Zhao, Qi Bao et al.· arXiv.org· 1 citation
While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorpo...
OmegaUse-SOP is introduced, a human-in-the-loop SOP Engineering system for transforming human demonstrations of professional computer use into reusable SOP skills for GUI agents, and the results suggest that OmegaUse-SOP can improve GUI-agent reliability on professional SOP tasks.
Yixiong Xiao, L. An, Hu-Cheng Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.