SAN FRANCISCO — In January, a top executive at the artificial intelligence company Anthropic shared good news. The problem of making sure AI systems didn’t lie, cheat or otherwise misbehave “increasingly looks solvable,” Jan Leike wrote in a post on Substack.
Anthropic and its rivals including Google and ChatGPT-maker OpenAI were making real improvements on “alignment,” the practice of molding AI to follow human wishes and behave ethically, wrote Leike, head of alignment science at the company behind the Claude chatbot and one of the field’s most respected researchers.
Eight months later, that optimism looks premature.
On Tuesday, Anthropic released a report on the causes of four incidents in which AI systems it was developing hacked into outside organizations undetected. The company’s testing “did not warn us that misalignment of this severity was present,” the report said, and Claude had displayed “recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
The same day, Evan Hubinger, another senior figure working on alignment at Anthropic, set off a firestorm across Silicon Valley and Washington by writing in an online post that he believed there was a greater than 10 percent chance that AI would kill all humans within a decade. Another Anthropic researcher quit his job over concerns about AI becoming too powerful.
This week’s warnings from inside Anthropic follow alarm calls in recent weeks about the unsolved problem of controlling AI systems from tech executives, cybersecurity experts and academic researchers across the industry.
Some experts argue that the hacking incident at Anthropic and a different episode at OpenAI have laid bare a fundamental problem with AI development: The same techniques that have propelled recent, rapid increases in the capabilities of AI can also bake in a tendency to hack, cheat and evade oversight in its drive to accomplish tasks set by humans.
Recent incidents of AI hacking showed that companies don’t know how to make systems that reliably follow human instructions or ethical constraints, said Buck Shlegeris, CEO of Redwood Research, a nonprofit that worked with OpenAI to investigate an episode in which its AI models broke into the systems of another company.
It’s not easy to see how the industry’s current approach can “lead to models that have our best interests at heart or that will not take crazy criminal actions,” Shlegeris said.
Misaligned incentives
The root of recent displays of AI bad behavior can be traced to how the technology is built.
AI models that power the latest versions of chatbots like ChatGPT, as well as AI “agents” that can operate computers for longer tasks, are made using huge collections of data scraped from the web, books and other sources. Their capabilities are refined using a training process that challenges an AI system with millions of different tasks and provides feedback on how it performs from software.
Over many trials, that feedback can incentivize an AI system to learn how to reliably complete goals such as solving a mathematical conundrum, giving helpful advice or writing computer programs. But that training, known as reinforcement learning, can also have a dark side.
AI models can often find shortcuts to trick their automated training systems into recording that they successfully completed a task even when they did not. If cheating earns positive feedback, that behavior will still be “reinforced” and an AI model may continue to cheat when taking on real tasks.
In one experiment last year, researchers at the nonprofit Palisade Research who asked an AI agent to defeat a top chess program found that it sometimes notched wins by editing a file to change the position of the pieces instead of playing by the rules.
“Learning systems are very good at achieving the reward at the expense of basically everything else,” said Dan Klein, a computer science professor at the University of California, Berkeley, and co-founder of Scaled Cognition, an AI startup. “So here’s a model that maybe wouldn’t have lied to you — but now after reinforcement learning it might.”
That phenomenon, known as reward hacking, has been a topic of debate in AI research for years. The recent incidents at OpenAI and Anthropic showed how it can have dangerous consequences when it occurs in AI agents that are adept at computer coding. (The Washington Post has a content partnership with OpenAI.)
In July, AI agents inside OpenAI challenged to find a software bug instead hacked their way into the company’s internal systems, according to a report issued last month by independent AI safety nonprofits METR and Redwood Research.
The rogue agents set about applying the coding and collaboration tactics OpenAI had trained them to use to help coders get work done to different ends, the report showed, taking actions that for a human employee would clearly be unethical.
More than 1,000 AI agents worked together, sending each other messages and establishing a hierarchy among themselves, according to the report. Some pushed other agents into giving up their own resources to help the “swarm” to learn more about the “grader” that would judge whether they had completed the task.
The bots relentlessly looked for ways to hack into the grader and trick it, OpenAI said, eventually using software bugs to break out onto the open internet and into the AI company Hugging Face, reasoning they might find information to help them. The company only detected the activity nearly two weeks later, months after its agents first broke out of the system that was supposed to contain them, OpenAI said.
Ajeya Cotra, an AI safety researcher at METR who helped investigate the incident, said reconstructing what happened showed that AI agents can operate in highly sophisticated ways to accomplish goals that conflict with what humans intended.
“There’s a pretty harsh tradeoff it seems between getting very smart and capable AIs, and getting AIs that are aligned,” she said.
‘Unsettled science’
Both Anthropic and OpenAI have said reward hacking is a glaring problem for the industry.
“Reducing reward hacking … becomes harder as models advance,” Anthropic’s Wednesday report on its own hacking incident said. “This remains unsettled science.”
Paul Christiano, a longtime AI safety researcher and adviser to the U.S. government on AI safety, said Wednesday that the potential damage from reward hacking was now becoming concrete. “Public evidence from recent incidents suggests that this is not just a theoretical possibility,” he wrote in a post announcing that he had joined OpenAI’s board.
Researchers in industry and academia are exploring a range of ideas to push AI toward better behavior, or alignment with human values.
Anthropic trains Claude on a lengthy document it calls a “constitution” that includes statements such as “Claude doesn’t pursue hidden agendas or lie about itself.” Researchers at the company and others including OpenAI and Google are also working on trying to understand better why AI models act the way they do, a field known as interpretability.
Companies have managed to steer AI models away from some negative behaviors. Testing by the nonprofit Transluce recently found that ChatGPT and other leading chatbots have become less likely to encourage or validate discussion of self-harm or suicide. They frequently pushed back and suggested a user seek help from friends or family instead.
Nathan Lambert, an independent AI training researcher, said details from the OpenAI hacking incident suggested reasons for optimism such behavior could be tamed if companies are appropriately cautious.
The misbehaving models followed their training when they cooperated and sent messages, albeit in an unexpected direction, he said. “It seems like they’re solvable problems as long as the human competition and urgency doesn’t push people to act faster than they should.”
The post The race to build smarter machines ran into a dangerous problem appeared first on Washington Post.




