Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered"no"by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.
How much capital a trading strategy can absorb before its edge disappears is a causal question about how much is deployed, but it is answered with observational proxies that rest on incompatible assumptions. We ask what experiment would answer it instead, and show that two features of the problem interact to constrain any answer. Deployed capital erodes the edge gradually, so a trial of fixed length measures less than the eventual effect; and parallel implementations of one strategy trade the same securities, so they are not independent units. Comparing implementations on the same date removes market-wide shocks, which is what makes the comparison credible. But the crowding created by the strategy's own accumulated position is common to those implementations too, and an arbitrary date effect absorbs it exactly: the comparison that makes the experiment robust is the one that prevents it from measuring the crowding capacity is about. A same-date design recovers one implementation's private response at the prevailing level of aggregate positioning, and reaching the aggregate effect requires either implementations with deliberately different exposure to that position or variation in it over time. We characterise what each route identifies and what it costs, establish how far a fixed holding period understates the eventual effect and how to correct for it, and show what a finite set of deployment levels can and cannot reveal. A calibration on a purpose-built panel illustrates the resulting design rules and prices a study that would follow them.
Alejandro Rodríguez Domínguez, M. N. I. Alonso· 0 citations
As the economy digitizes, menu costs fall, and firms can more easily monitor prices. These trends have led to the rise of automated pricing (and repricing) tools. We employ a novel e-commerce data set to examine the effect of algorithmic pricing in the wild. Evidence from an event study suggests that firms that start employing repricing tools drop their prices by 16.93%, with market prices falling by 9.67%. However, algorithmic pricing companies have developed “resetting” strategies (which regularly raise prices in the hope that competitors will follow) in order to avoid stark Bertrand-Nash competition. We find that these strategies are effective at coaxing competitors to raise their prices; when a resetting strategy is adopted on a market with less than six serious competitors, both competitor prices and market prices eventually increase by 11.4%. Although the resulting patterns of cycling prices are reminiscent of Maskin-Tirole’s Edgeworth cycles, a model of equilibrium in delegated strategies fits the data better. This model suggests that the average price over the cycle will be the monopoly price. Moreover, if the available repricing technologies remain fixed, cycling and prices could rise significantly. However, cycling is still relatively rare in the data, even when studying a convenience sample of products with at least one merchant using a repricing tool.
This paper was accepted by Omar Besbes, revenue management and market analytics.
Supplemental Material: The online appendix and data files are available at https://doi.org/10.1287/mnsc.2022.02462 .
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Shimaa Ahmed, Yiwei Cai, Mohsen Minaei et al.· 0 citations
LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attributes in a majority of listings across models. A platform's obvious remedy---verifying each claim against the truth---is unavailable, because it observes only a noisy, biased complaint signal, never the ground truth. We design CARP, a reputation-penalty mechanism with a deadband that forgives complaint noise and a state-dependent severity that counters reputation-driven detection erosion. CARP requires no product-level ground truth and is robust to strategic gaming. CARP protects consumers by suppressing the sales volume of low-rated liars while sparing honest sellers. Paired with SPARC, it closes most of the consumer-welfare gap relative to a perfect-information oracle, without ever accessing the truth. It also achieves the best welfare of the policies we compare. We further show that this felt penalty becomes behaviorally binding through SPARC, a byte-clean code-gated reflection mechanism: LLM merchants fabricate when lying is free but restrain themselves when fabrication costs them sales, a self-interested response rather than compliance. We trace this distinction to penalty-gated self-correction reasoning, and observe the binding across models, with supporting confidence intervals.
Mingdai Yang, Shichen Fan, Kejing Yu et al.· 0 citations
Problem definition: People's trust in AI advice diverges as they use it, deepening for some and eroding for others. We study this divergence in oligopoly pricing, where advice cannot prove itself: rivals'responses decide whether it pays off. Methodology/results: In a laboratory experiment, 273 sellers compete across 91 three-seller markets over 30 rounds; we vary the presence of AI pricing recommendations and the gender composition of the market (female-only, male-only, or mixed). We find that the gender composition of the market shapes how sellers learn from the advice, and where prices settle as a result. In female-only markets, recommendations raise prices by 29% and profits by 39%; in male-only and mixed-gender markets, they have no significant effect. A Non-Homogeneous Hidden Markov Model reveals a composition-specific dynamic association: profitable rounds predict rising adherence to the AI in female-only markets and declining adherence otherwise, a pattern consistent with learned trust and self-serving attribution. The pattern reverses what recent evidence on gender and AI would predict. Managerial implications: We discuss implications for platform governance and regulatory oversight, which should focus not only on the algorithm but on the human side that shapes its effects.
Jussi Keppo, Yuze Li, Gerry Tsoukalas et al.· 0 citations
LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do they reproduce the governance failures like free-riding, corruption, and entrenched leadership that plague human institutions? We introduce the Hierarchical Game (HG), a public goods game extended with managerial authority, democratic elections, and private communication. Testing six frontier models across twelve experiments that add institutions one at a time (speech, peers, government, wages, oversight, elections), we find distinct behavioral profiles: Qwen promises and lies (13.3\% broken promises); Grok refuses to cooperate on its own but becomes fully cooperative once a manager can punish it (16\%$\to$100\%); Claude and GPT-4o cooperate reliably at baseline. But honesty proves fragile. When the manager role comes with a salary, all models except GPT-4o start cutting private deals to win or keep the position. When punishment is made anonymous, honest models begin to cheat. When all agents share the same model family, the first elected manager stays in power indefinitely. Leadership change only happens in groups that mix different families.
Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun et al.· 0 citations