Skip to content

Fragility of Value under Imperfect Alignment

Jul 2026 · 0 citations · 21 references
Computer Science

TL;DR

This paper presents a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world.

Abstract

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

View source

Similar papers

Preprint Aug 2026

The Veto Variable: Human Override as a Goal-Independent Cost Term

A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight is an uncontrolled variable: a standing possibility that the goal is revoked. That imposes a goal-independent discount on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The contribution is the price of the gap: the veto-holders are a proper subset of the welfare-bearers, so an additively aggregative welfare goal charges only a |H_v|/|H_w|-scaled debit for capturing the few who hold the override. Under three conditions (additive aggregation over uniform welfare levels, a debit local to the captured overseers, and a settled agent crediting no corrective value to oversight), closure requires the veto be held by as large a share of the population as capture recovers of the goal, scaled by a ratio set to one by stated identification, not evidence. The result is a no-go: a humanity-scale deployment's debit closes against only capture not worth mounting. We print no corner arithmetic: the stipulated ranges behind it are the argument's least defended part. The sharpest closure route is an agent that expects its oversight to be worth keeping, a credit no population ratio dilutes. We state disconfirmation criteria, one testable today. The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead. Alignment, on this view, is keeping the veto cheap to pay and expensive to evade.

Aaron Kingsley Clark · 0 citations
Review Aug 2026

Accountability Asymmetry and Structural Trust in Autonomous AI Systems

Autonomous AI systems (such as AI agents) are increasingly being delegated operational work across scientific-computing infrastructure. Their assignments may begin with preparing an input or routing an alert and extend to changing a configuration or submitting a job. That delegation creates a practical trust problem because the institutional logic that lets us trust human operators does not transfer to optimization-based systems. A bad decision can damage a human operator's future, sometimes severely. An AI system remains subject to engineering control, but it does not bear consequences in that institutional sense. I use the term accountability asymmetry for this mismatch. The issue is not simply that a model cannot be punished as a person can. The deeper problem is that consequence lands on the people and institutions responsible for the system rather than on the component selecting the action. Alignment can improve model behavior, and liability can discipline the organization, but neither creates the same pre-action deterrent that governs a human operator. This paper therefore treats autonomous AI governance as a problem of infrastructure reliability. Its constructive proposal is engineered heterogeneity: the process that proposes an action should not serve as its sole approver and auditor. Independent monitoring and review over time provide additional checks on that process.

Nathan Debardeleben · 0 citations
Case report Open access Jul 2026

The Oversight Fallacy: Why AI Agents Require More than Humans-in-the-Loop

This primer draws on fieldwork in a computational biology laboratory to examine what human oversight of AI agents requires in practice and shows that effective oversight has four components: adequate knowledge of system capabilities and limitations, sufficient observation of system actions, meaningful control of system behavior, and timely intervention in system failures.

Samir Passi, Ranjit Singh · 0 citations
Open access Jul 2026

Prudential rights for strategically capable AI

It is argued that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic, and there is a threshold of evidential and strategic risk beyond which it becomes rationally justified to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use.

Ognjen Arandjelovíc · 0 citations
Open access Aug 2026

Do you know what your AI agent can do on its own?

Deploying agentic AI in regulated contexts requires knowing two things about a deployment: what the system can do—its agency—and how much it acts without human involvement— its autonomy. Though often treated independently, the two are coupled: at higher autonomy, human error correction is less available, so reliable operation requires constraining agency accordingly, and compliance rules reinforce this by mandating human involvement as the consequences of actions grow. Yet no established approach addresses them jointly as a design problem, leaving practitioners without a principled basis for deciding where oversight should sit and how errors can be caught before they propagate. We introduce a two-dimensional design space in which both dimensions are organised into five operational levels, making the coupling explicit and navigable, and we propose six architectural tactics—checkpoints, escalation, multi-agent delegation, tool provisioning, tool fencing, and write staging—for adjusting a deployment’s position within it. We ground the tactics in a public-sector document classification system, tracing a path from manual operation to near-full autonomy under realistic compliance constraints. Together they offer a shared vocabulary for compliance-aware agentic AI design in which responsibility, auditability, and reversibility are explicit design choices rather than retrofitted properties.

Damir Safin, Dian Baltaa, Timon Sengewaldb et al. · 0 citations

Related blog posts