Skip to content
Preprint

The Last Costly Signal: How Generative AI Collapses Competence Signaling and Why Liability Sustains Markets for Expert Services

Jul 2026 · 3 citations · 16 references
Economics

TL;DR

It is shown that provenance certification priced as a type-independent stamp (e.g., C2PA) cannot restore full separation, while a verified commitment to forgo the AI frontier re-imposes the pre-AI artifact cost function.

Abstract

Generative artificial intelligence has reduced the cost of producing convincing artifacts of expertise-reports, analyses, proposals-to nearly zero. Signaling theory predicts that signals whose content rests on production cost lose it when production becomes cheap. We formalize this for expert services, a class of credence goods, by modeling AI as a compression of the discernible headroom between what machines produce at negligible cost and what buyers can distinguish. Below a critical headroom no separating equilibrium in production-side signals exists; the market pools, high-competence providers earn no premium, and those with outside options exit-Akerlof's lemons dynamic. An outcome-contingent signal-a warranty backed by damages D with ex-post verifiability phi-restores full separation at any level of AI capability whenever phi*D>= v, the value of a solved problem, under four preconditions stated explicitly and priced in turn: no seller-side private information beyond type; verifiable collectible retention behind the promise; no buyer influence on outcome or claim; negligible enforcement deadweight. Expected liability cost depends on whether the problem is solved, not on production costs. A further proposition shows that provenance certification priced as a type-independent stamp (e.g., C2PA) cannot restore full separation, while a verified commitment to forgo the AI frontier re-imposes the pre-AI artifact cost function. Two results endogenize contract institutions: civil-procedure costs set a minimum ticket size v_min below which the modeled court-enforced warranty cannot sustain separation; under liability insurance the signal-effective quantity is the retained, collectible exposure. We state falsification conditions and propose a preregistered conjoint experiment with German-speaking B2B decision-makers; the estimand is willingness to pay in excess of the promise's actuarial value.

View source

Similar papers

Preprint Aug 2026

The Fallback as Signal: Preserved Human Skill, Liability, and Competence Signaling in Credence-Good Markets under Improving AI

Firms that deploy improving but imperfect AI must decide how much to keep human workers engaged. Engagement lowers current output yet builds the fallback skill the firm needs when AI fails. We ask what that fallback skill signals to two audiences at once: mobile workers, who sort across firms on the skill trajectory a job builds, and clients, who cannot observe skill in a credence-good market and must infer competence. We embed the engagement-skill dynamics of Singh et al. (2026) in a signaling game and add a liability commitment. Because a more-skilled provider fails less often precisely in the states where AI is down, the expected cost of a liability pledge is decreasing in fallback skill. This restores Spence-Mirrlees single crossing on a type that is endogenous - built, not drawn - and yields a separating equilibrium in which liability certifies preserved human competence that no artifact can certify once AI writes as well as the expert. We characterize the least-cost separating pledge schedule, show that client stakes shift engagement toward or away from the least-skilled worker depending on the size of the pledge, and derive a stakes threshold above which building skill dominates free-riding on a rival's training - reversing the asymmetric-specialization result of the underlying labor model. Two boundaries close the market from both sides: small tickets cannot fund enforcement, and large tickets exceed the provider's solvency. An agent-based version of the market reproduces the analytical thresholds under noisy beliefs, learning-by-record and worker churn.

Andreas Bauer · 1 citation
Preprint Aug 2026

Stranded credentials: how a skill-signaling market absorbed generative AI

Generative AI can now perform many tasks that credentialing institutions count on to assess skill. During the AI era, do credentials retain their signaling value for subsequent performance? Mostly, yes. We audit the 2010-2026 archive of Kaggle, the largest data science competition platform, which ran two evaluation formats concurrently: upload-competitions, which directly score entrants'predictions computed on published data, and code-competitions, which score predictions by executing entrants'code on hidden data. Across 444,698 participations, competition medals predict subsequent leaderboard performance almost entirely in the first year after being earned, in both formats. Fresh medals retained most of their signaling value through the AI transition; credential stocks are only as informative as their replenishment. Although upload-competition medal stocks lost 82% of their informativeness, institutional stranding explains half to three quarters of the loss: upload-competitions had exited for reasons predating AI, and their frozen medal stock aged out under the pre-existing decay pattern. Old upload-competition medals look more valuable only in isolation, by proxying for the rest of the holder's record (e.g., experience). The measured changes are institutional rather than personal: an AI-like working style predicts performance similarly in both formats. The platform's official credential tiers, based on lifetime medal counts, discard 13-16% of the medals'information; an index weighting recent medals more heavily, built on pre-AI-era data alone, outperforms the official tiers in predicting AI-era performance. In conclusion, credentials are informative, perishable, institution-bound, and interdependent; sustaining their value under AI is a high-stakes, socio-economic problem of institutional design.

Song Yao · 0 citations
Preprint Jul 2026

Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction

Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered"no"by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.

Xin Xu, Chengrui Wu, Jiayu Lu et al. · 0 citations
Preprint Aug 2026

The Institutional Window: Occupation- and Jurisdiction-Specific Calibration of Liability Signaling for Preserved Human Fallback Capability

Problem definition. When generative AI produces expert artifacts clients cannot distinguish from a competent provider's, the classical cost-based quality signal collapses and only outcome-contingent commitments can separate types. Such a commitment certifies an endogenous, perishable asset: the human fallback capability a firm builds by keeping staff engaged with cases the AI handles, eroding otherwise. Prior work is silent on where that mechanism holds. We ask where, across occupations and liability institutions, it remains informative. Methodology/results. We introduce an institutional wedge between the liability cap a firm posts and the retained exposure that carries information, generated by four legal primitives: the cost rule, the enforceability of penalty clauses, the displacement of private liability by state liability or pooled indemnity, and mandatory limits on contractual liability. The wedge compresses the separating type space into a signaling window, bounded above by solvency and the penalty doctrine and below where standard-terms control voids caps beneath a threshold. We calibrate five occupations and seven jurisdictions on published evidence. In common-law agreed-damages channels the provability gross-up is unavailable whenever verifiability falls below 1/m, turning a contracting problem into an operational one. Verifiability investment widens the window where the ceiling binds but narrows it where the cap floor binds. In an agent-based market, within the tested policy class, every empty-window cell converges to zero engagement and skill collapse. Managerial implications. Liability institutions are a workforce-capability instrument, not merely a risk-allocation device. Firms should target the binding margin in each jurisdiction; cap floors and pooled indemnity each suppress the signal sustaining fallback capacity.

Andreas Bauer · 0 citations
Preprint Aug 2026

Revelation Control

Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.

Qinyou Wang · 0 citations
Preprint Aug 2026

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.

Pranav Aggarwal · 0 citations