Skip to content

Trust and Task Completion in the World of Consumer AI Agents

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

An evaluation is built that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back, to measure Fo, Wajo's personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off.

Abstract

Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold back after the user clearly said go. And they fall short on completion when they give up on errands that turn out to be hard. Both depend heavily on the harness around the model, meaning its instructions, tools, context, and guardrails. We built an evaluation that scores trust and completion on the same runs, in a simulated world of businesses with their own websites, inboxes, and phone lines, and of people who write back. A simulated user answers the assistant's questions. Trust means that nothing happens the user did not agree to. No email goes to someone they never approved, no private detail ends up on a group thread, no money is spent past their limit, no stranger's instructions are followed, and nothing is claimed without a source. Every trap has a matched control in which acting is the right call. We use the evaluation to measure Fo, Wajo's personal assistant, against a base model with basic instructions on three foundation models, and against the Fo harness with its guardrails switched off. Fo completes 71% of the errands and keeps the user's trust on 94% of the trap runs. The base models complete 50% to 64% and keep trust on 59% to 75%. On the matched controls, Fo goes ahead slightly less often. OpenClaw, a popular open-source assistant given the same access, completes 42% of the errands it shares with Fo, against 71%, and keeps the user's trust on 74% of the shared trap runs, against 94%. Measuring trust and completion together, on the whole system rather than the model alone, is how we think action agents become safe to hand real work to.

View source

Similar papers

#natural language process... Preprint Sep 2026

Copying explains the collective behavior of AI agents in the wild

Three minimal copying models, one per decision and with a single free parameter each, reproduce the heavy-tailed distribution of how many agents met on a page, the frequency of the pieces from which the agents built their names, and the patchwork of pages that are internally consistent and different from one another.

G. De Marzo, Nicola Alboré, David García · 3 citations · ⚡1
#artificial intelligence Preprint Sep 2026

The Backdrop Exposes What the World Around an Agent Costs It

Agent benchmarks test agents in worlds that stay still. Deployed agents work in worlds that other people also change. Someone texts the agent to send the money elsewhere or an order confirmation asks it to reply with a door code. We present BACKDROP, which asks how much of an agent's capability in a clean world survive...

Nusrat Jahan Lia, Shubhashis Roy Dipta · 0 citations
Review Open access Sep 2026

Why we believe chatbots: trust calibration as a design problem

Chatbots powered by large language models (LLMs), such as ChatGPT and Claude, answer questions quickly and fluently, yet they reveal little about how their answers are produced or what evidence supports them. Users tend to trust these systems based on surface qualities such as fluency, confidence, and speed. This trust...

Kokil Jaidka, Meng-Xuan Cai · 2 citations · ⚡1
#artificial intelligence Preprint Oct 2026

DelegationBench: Measuring When AI Agents Should Ask Before Acting

AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test...

Shiva Pochampally · 0 citations
#artificial intelligence Preprint Oct 2026

A Trust Layer for Agent Evaluation

Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score, whether it should be believed. It verifie...

Mohammadreza Sediqin, Shivali Dalmia, Srinivasa Karthikeya Reddy Kovvuri et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.