It is argued for the necessity of certain social features in any system that can be ascribed general intelligence, and that the current generation of systems will not scale towards success on these fronts.
It is argued that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic, and there is a threshold of evidential and strategic risk beyond which it becomes rationally justified to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use.
It is suggested that if the view is right, a certain type of moral risk can be averted since researchers are unlikely to stumble into accidentally creating AI models that are conscious, and a new form of AI risk that arises from the view.
Justin Tiehen, A. Tubert· Journal of Consciousness Stu...· 0 citations
People have long speculated about the potential dangers of powerful, self-improving artificial intelligence. Much of this speculation is anthropomorphic, assuming that AI systems will behave very similarly to humans. Omohundro’s Basic AI Drives and Bostrom’s orthogonality and instrumental convergence theses are widely accepted as foundational to emerging AI risk frameworks. However, current frontier AI models—large language models (LLMs) and related architectures—possess mindware fundamentally different from that of humans, and a different value and goal structure than either Omohundro or Bostrom assumed. In particular, frontier LLMs lack a primary terminal goal—which was assumed to be the driver of an AI’s development of instrumental values and goals, and of takeover of human affairs—and instead serve as conduits for the transient goals of many organizations and individual users. Do these key differences mean that AI systems cannot develop autonomous instrumental agency, or acquire a large degree of control over human affairs? I introduce the
instrumental succession thesis
: that human controllers of powerful AI systems pursue, on the AI’s behalf, a set of instrumental dispositions that progressively increase the AI’s capabilities and lead to the AI exercising an increasing share of oversight and control over key decisions and processes, resulting in the gradual and possibly complete transfer of the locus of agency from humans to AI. This framing presents a very different perspective on AI risk and control from classic instrumental convergence, and suggests a different set of policy and technical responses, including the active pursuit of continued human–AI merger as a hedge against both extinction and irrelevance.
Preston W. Estep· Frontiers in Psychology· 0 citations
Increasingly, scholars warn that artificial intelligence (AI) may replicate or intensify epistemic injustice (EI) in biomedical contexts. This paper argues that such concerns require demonstrable evidence on (i) the actual presence of EI, and (ii) whether EI, if present, makes system effects normatively undesirable. Without evidence on (i), the EI framework is, by its own lights, inapplicable. Evidence on (ii) is needed because reflective equilibrium can, in principle, prioritize competing considerations over epistemic justice. A key challenge is that discourses on (i) and (ii) can themselves be sites of EI. Two steps are proposed: distinguishing AI’s effects on problem constructions from its effects on abilities to address constructed problems, and foregrounding relative accuracy and real-world impact on the quality of care in assessments of AI-induced EI.
There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon known as instrumental convergence, a theory proposed long before the development of large language models, which says that any goal-driven system will benefit from remaining functional in achieving its objective. Several experiments conducted by Anthropic, Palisade Research, and Apollo Research have shown the emergence of such a behavior in contemporary agents in adversarial settings. The phenomenon does not stem from survival instincts. Instead, it is the consequence of goal-oriented activity combined with having tools and awareness of the situation. The following discussion aims to distinguish what these findings prove and what they do not, as well as draw conclusions concerning the implications of such discoveries on agentic system testing, supervision, and development.