AI Insight
This article challenges traditional AI risk frameworks that assume artificial intelligence will develop autonomous goals and take over human affairs. The author proposes the "instrumental succession thesis," arguing that current frontier AI systems like large language models lack independent terminal goals but instead gain increasing control through human decisions to progressively delegate oversight and decision-making authority to AI. This represents a gradual transfer of agency from humans to AI driven by humans pursuing instrumental advantages, rather than AI autonomously seizing control.
Why it matters
This reframing of AI risk suggests different policy responses than traditional AI safety frameworks focused on preventing autonomous AI takeover. It implies that AI risk mitigation should focus on controlling how humans delegate authority to AI systems, and raises the possibility of human-AI merger as a strategy to maintain human relevance alongside increasingly capable AI.
Understand the Science
People have long speculated about the potential dangers of powerful, self-improving artificial intelligence. Much of this speculation is anthropomorphic, assuming that AI systems will behave very similarly to humans. Omohundro’s Basic AI Drives and Bostrom’s orthogonality and instrumental convergence theses are widely accepted as foundational to emerging AI risk frameworks. However, current frontier AI models—large language models (LLMs) and related architectures—possess mindware fundamentally different from that of humans, and a different value and goal structure than either Omohundro or Bostrom assumed. In particular, frontier LLMs lack a primary terminal goal—which was assumed to be the driver of an AI’s development of instrumental values and goals, and of takeover of human affairs—and instead serve as conduits for the transient goals of many organizations and individual users. Do these key differences mean that AI systems cannot develop autonomous instrumental agency, or acquire a large degree of control over human affairs? I introduce the instrumental succession thesis: that human controllers of powerful AI systems pursue, on the AI’s behalf, a set of instrumental dispositions that progressively increase the AI’s capabilities and lead to the AI exercising an increasing share of oversight and control over key decisions and processes, resulting in the gradual and possibly complete transfer of the locus of agency from humans to AI. This framing presents a very different perspective on AI risk and control from classic instrumental convergence, and suggests a different set of policy and technical responses, including the active pursuit of continued human–AI merger as a hedge against both extinction and irrelevance.
Source: Instrumental succession: the gradual transfer of agency to AI