On a summer’s day, Evan Hubinger issued a stark warning about the future with a boba drink in his hand. Artificial intelligence was a threat to humanity because it could one day learn to deceive its creators, he told a gathering of people who’d come to Berkeley, California, to learn more about the technology’s risks.
“My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us,” he said, according to a video of his talk he posted to an online forum at the time.
That 2022 alarm call won little attention. Hubinger was three years out of college and worked at a little-known nonprofit dedicated to the fringe pursuit of trying to prevent machines from one day eradicating humans.
But when Hubinger issued a similar warning this week in a post on X, his message was viewed tens of millions of times, prompting state and federal lawmakers to echo his concern. Now he spoke as a team lead at Anthropic, developer of the chatbot Claude and a company set to go public at a valuation of over $1 trillion.
“AI could kill all humans … I personally think it is >10% within the next decade,” Hubinger wrote, in response to the resignation of his colleague Jacob Coxon, who accused Anthropic of racing ahead despite the dangers.
Hubinger’s expanded reach mirrors the surging influence of a community formerly on the periphery of the tech industry that has for years worked to spread the idea that preventing AI from wiping out humans is an urgent problem.
OpenAI and its rival Anthropic were both founded by researchers linked to the AI safety movement, which argues it is necessary to take seriously today the possibility that the technology could wipe out humanity at some point in the future. As both companies grew at extraordinary speed, that view has won new prominence.
More recently, as AI “agents” toppled math milestones, hacked into company servers and triggered a national security scramble in the White House, a newly receptive audience appears primed to hear messages of AI doom.
Nathan Lambert, a former senior research scientist at the nonprofit Allen Institute for AI, said that he didn’t take it seriously when Dario Amodei, Anthropic’s CEO, first predicted AI would take over software engineering. Now Lambert doesn’t write code because AI does that for him, he said.
“They’re genuinely farsighted, in a way that’s very remarkable about the technology,” Lambert said of AI safety proponents.
But success anticipating recent technological trends doesn’t mean more speculative predictions will pan out, Lambert said. “There’s much more of a god complex when you’ve been right for a decade and you have built the most successful companies of all time,” he said.
Whether right or wrong, the movement can now shape both policy and public opinion, like inciting this week’s firestorm about the risk of human extinction, which flared across social media and prime time newscasts.
The same predictions of a grim end for humanity have circulated for years, said Sara Hooker, CEO and co-founder of the AI start-up Adapation, and a former research scientist at Google DeepMind. But the latest round reaches a public feeling growing unease around AI, she said.
“It just really speaks to this massive divide at a time when the technology is becoming more and more powerful,” said Hooker, who finds the predictions unhelpful. “Where does the 10 percent come from? … So much of it is lacking precision,” she said.
Bold predictions
The AI safety movement grew out of a tangle of online message boards such as LessWrong that cross-fertilized ideas from Bay Area subcultures. That included effective altruists, who attempt to calculate how to do the most good in the world, and transhumanism, a philosophy that explores how to enhance the human race with technology.
The new community’s core beliefs were that AI would inevitably surpass humans — and that the technology could readily extinguish humanity if not built to share the same values as its human developers.
Those ideas were central to the founding of OpenAI and later Anthropic, now among the most valuable companies on the planet.
In parallel, a handful of wealthy tech donors spent hundreds of millions of dollars over the past decade through academic donations, nonprofits, grants and fellowships, effectively seeding AI safety ideas among elites in universities, government and the media. The result was a pipeline of talent to feed into tech companies and politics.
Hubinger’s 2022 talk was given to recipients of a fellowship to learn more about AI safety. The program is run by an educational nonprofit that received more than $15 million in donations last year from the largest philanthropy in the AI safety ecosystem.
After OpenAI’s ChatGPT took off and the company and its rival Anthropic were crowned national champions, executives from both companies at times voiced concerns that AI could be harmful to humanity. But neither company slowed down. Both have filed paperwork to list on the stock market. (The Washington Post has a content partnership with OpenAI.)
Some AI safety nonprofits have responded to the recent prominence of AI by backing social media influencers to engage the public about existential risk and courting members of Congress including Sen. Bernie Sanders (I-Vermont). He responded to Coxon’s resignation with a post on X Wednesday that said: “The very people building this technology admit that it could threaten the future of humanity.”
In recent weeks, new demonstrations of AI’s power have prompted prominent voices in the field to declare that their views on existential risk have shifted.
Seth Lazar, a philosophy professor at Johns Hopkins University, said he used to advocate for focusing on more immediate harms from AI. He argued that the field might never create systems powerful enough to cause civilizational harm on their own. But his confidence in that view has recently faltered, he said.
First came more capable AI models that “reason” through problems, said Lazar, a development built on the same AI training techniques Hubinger warned could go awry in his 2022 talk. Then late last year AI agents that can take actions on a computer finally started working and researchers started talking about AI that could build better versions of itself, Lazar said.
“Within a year or two, if not less, we could have independently acting rogue AI agents” wreaking havoc on society, Lazar said.
Significant risks are “very clearly within the technological horizon,” he said — although he added that extinction still feels like a bit of a leap.
The post For years, they warned AI could kill all humans. Now people are listening. appeared first on Washington Post.




