Skip to content

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

This work identifies a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable.

Abstract

Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.

View source

Similar papers

Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...

Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al. · 1 citation
Preprint Aug 2026

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is...

Pranav Aggarwal · 0 citations
Preprint Aug 2026

Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring

It is well established that rendering documents as structural markup rather than flat prose improves extraction, and it is well established that rendering documents as structural markup rather than flat prose improves extraction, and that on three reading tasks is reproduced.

Cheng Yu, Nikhil Mathew, Zhengjie Wang · 0 citations
#artificial intelligence Preprint Sep 2026

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks, and contributes a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting.

Pei-Ke Zhu, Si-Di Chang · 2 citations
Open access Sep 2026

Transparent but Trustworthy: Reconciling the AI Disclosure Paradox in Corporate and Customer Communication

Regulators and public interest activists increasingly advocate for full transparency in customer-facing AI. However, an empirical meta-synthesis of 17 studies (N=14,820) reveals a critical transparency dilemma: uncontextualized raw AI disclosures trigger negative persuasion knowledge, reducing consumer trust by 1.12 to...

Sana Hussan · 0 citations
Preprint Aug 2026

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

This work argues that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure, and introduces DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymm...

Neel Tushar Shah, Manglam Kartik, Akshat Karkar · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.