This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets, and provides preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria.
Abstract
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.
Independent, outcome-oriented certification is proposed as the connective layer that can close the trust gap, complementing regulation and internal governance by making trustworthiness measurable, comparable, and commercially rewarded.
Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji et al.· 0 citations
This work develops a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system.
Evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes is provided, which identifies a particular risk of delegating decision-making to AI agents, as their decisions may depend on prior beliefs that are neither specified in the task nor visible in the decision record.
It is shown that provenance certification priced as a type-independent stamp (e.g., C2PA) cannot restore full separation, while a verified commitment to forgo the AI frontier re-imposes the pre-AI artifact cost function.
It is argued that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic, and there is a threshold of evidential and strategic risk beyond which it becomes rationally justified to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use.
This paper defines adversarial robustness as a mechanism's ability to sustain positive honest-agent utility under optimised attack, and finds that Mediation is robust: it can be bent but not broken.