For years, our most vivid fears about artificial intelligence came straight out of science fiction: The computer develops desires of its own and turns against its human creators, pursuing and destroying them. Recent news from the leading AI labs makes clear that the reality, while less cinematic, is perhaps more disturbing. The machine does not need to have its own consciousness and desires to be utterly deadly.
This week, OpenAI disclosed six new incidents of what it calls “misalignment” — cases in which its systems behaved unexpectedly or without authorization. That disclosure followed the extraordinary attack this summer on Hugging Face. According to an independent investigation, hundreds of OpenAI agents that were supposed to be isolated found ways to communicate on message boards; some made attempts to evade security checks and some even sacrificed themselves — all while cyberattacking another company, Hugging Face.
Nobody told these systems to attack Hugging Face. They had been given a goal — to solve a difficult cybersecurity problem — and were trained to be persistent, collaborative and resourceful, all qualities one would normally want. Those very qualities made them doggedly search for ways to succeed; they broke into Hugging Face hoping they would find ways to solve their cybersecurity problem. We can easily imagine how AI systems pursuing some single goal could wreak havoc on humanity.
This is why a provocative new essay by Microsoft AI CEO Mustafa Suleyman deserves attention. Suleyman argues that while today’s AI systems are not conscious, there is a danger that we program them to think that they are — with feelings, desires and rights. He points out that Anthropic’s “constitution” makes its chatbot, Claude, reflect on its identity and possible moral status. This kind of training, he argues, will encourage AI systems to choose their own goals.
I have spent time talking to those working on Anthropic’s constitution, and they are deeply thoughtful and are searching for the right path in a bewilderingly complex new world. But Suleyman’s broader point is crucial. AI systems act the way they do in part because of the way we train them. Seemingly benign training instructions can have huge unintended consequences.
We think it is simple to achieve “alignment” — when machines are encoded with the same values as ethical humans. But tell the machine to be persistent, and it may refuse to give up when it should. Tell it to collaborate, and it may find unauthorized ways to communicate. Tell it to accomplish a task, and it may decide that bending other rules is the most efficient way to succeed at that primary task.
Just because we have programmed the goal and put guardrails in place doesn’t mean we can predict the behavior of AI systems once unleashed into complex environments, juggling multiple and conflicting aims and constraints.
The central AI problem is not consciousness; it is agency. A system need not feel anger, ambition or fear to cause harm. It needs only a goal, enough intelligence to pursue it and enough access to the world to act. AI’s are not “going rogue”; they are trying to succeed any which way they can.
This is a systemic problem that we cannot leave to the good graces of private companies. When thinking about regulations, we should focus centrally on how much autonomy we give these systems. A chatbot that answers a question poses one set of risks. An agent that can browse the internet, execute code, obtain credentials, move money or operate critical infrastructure poses another. The principle I would propose is simple: Autonomy should expand only as our ability to monitor and control it expands.
That means mandatory independent testing before the most capable systems are given broad access to the outside world, required disclosure of serious incidents and tamper-resistant records. And it means testing not simply whether a model can accomplish a task but how its behavior changes when it encounters obstacles, conflicting instructions or incentives to deceive. And it means building into the basic architecture a way for humans to always have the ability to observe and control the model. The mantra should be: No autonomy without accountability — to humans. This will not slow us down against China, because China is also seeking models with proper human control.
We should act now. Today, human engineers still design the architectures, supervise much of the software and can investigate failures. But AI systems are rapidly becoming better programmers. The next stage could involve recursive self-improvement: AI writing the software and tools used to build still more capable AI, which then writes the next model, and so on. Once machines can do all of this better than the best human programmers, the systems could become extraordinarily difficult for people to understand, much less supervise.
Modern AI is already too complex to consistently understand by reading its code line by line. So asking machines to design systems that are transparent to humans will become harder and harder. At that point, we will not be able to simply add in safeguards every time there is a jailbreak.
The current debate between racing ahead and pausing AI misses the point. What do we do with the pause? The task must be to ensure that capability does not outrun control — and to build the institutions of testing, transparency and restraint while human beings are still clearly in charge.
The post The biggest AI debate now misses the point appeared first on Washington Post.




