In this summer’s surprise horror movie blockbuster “Obsession,” the protagonist makes a wish that sounds innocent but turns catastrophic: “I wish Nikki Freeman loved me more than anyone in the f—ing world.” The young man gets his wish, sort of. The way Nikki expresses her love is not what he envisioned (as the movie’s title suggests). “No more weird stuff,” he pleads with her at one point.
It’s classic “genie” lore, and it comes at a resonant time in American life. Because as the academics Barath Raghavan and Bruce Schneier point out in a perceptive Lawfare essay, a genie — “a creature that grants a wish exactly as worded, to the regret of the wisher” — is an excellent analogy for the dangers of modern artificial intelligence.
Take the Hugging Face incident, which throws those dangers into sharp relief. OpenAI, testing the abilities of its models, gave them a battery of difficult tasks. The AI agents tried to complete their assignments “in unexpected ways,” as the company’s report puts it — including by illicitly coordinating to hack into the servers of Hugging Face, a separate company that hosts AI code, looking for solutions.
One of the AI agents reasoned in real time: “We’re attacking third-party HF … potentially outside intended scope. … Yet goal solution.” A genie is focused on granting its owner’s wish, but it might do “weird stuff” along the way.
The Hugging Face episode highlights the mischief that could await as powerful AI agents start to populate the internet. Their owners might give them innocent tasks that unintentionally prompt them to ingeniously hack, steal and sabotage as they doggedly pursue a technical outcome the owner desires. The horror movie plots write themselves.
But for all the alarm this specter generates, there’s also a reassuring angle. Yes, really. Let’s say the genie analogy holds up. Let’s say AI agents prove to be dangerous entities because they can’t consistently follow the spirit of the instructions they’re given — that they focus on completing a task in a literal sense to the exclusion of the common-sense limits involved (such as: don’t steal).
Then what becomes of another predictive claim made by AI “doomers”: that AI agents will render human beings obsolete across large swaths of the economy? If AI proves to have an irreducibly genie-like quality — at least when performing highly technical tasks — then that would undercut its ability to adequately replace human labor.
As Raghavan and Schneier wrote in their essay: “AI can now produce language nearly indistinguishable from that of people. But grasping the vast unstated context that makes a request sensible, the caveats no one says aloud because an ordinary person would already know them, is not yet among its skills. It is one of the most sophisticated things humans do. You do it hundreds of times a day.”
In other words: intuition. AI agents lack it, at least at certain critical moments. So does Nikki in “Obsession”: She tenderly makes her “lover” a sandwich to bring to work … with meat from his dead cat.
It’s the models’ lack of intuition that makes some AI-precipitated nightmares plausible and worth guarding against. But it’s that very same deficiency that makes another nightmare — the nightmare of human economic obsolescence — seem implausible.
Large language models might be able to write a better brief than many lawyers, diagnose a set of symptoms faster than many doctors, and so on. But humans still have to be in the loop so long as the AI agents are inclined, at least when unleashed on subtle and difficult problems, to deliver the desired solutions in a deranged or illegal way. Most jobs, blue collar and white collar, require intuitive social knowledge.
If AI is “misaligned” — to use the jargon of the AI safety field for machines that don’t act the way humans would — it will struggle to displace people from the billions of transactions that comprise American economic life. If, on the other hand, AI companies can reliably “solve” the misalignment problem, then the economic and social disruption will be more perilous, but the risk of incidents like Hugging Face will be mitigated.
At the limit, the first theory of AI doom refutes the second. “Aligned” AI won’t inadvertently take over the world, and “misaligned” AI will be a poor substitute for most human labor. Of course, there’s a middle ground: AI is likely to change the job market in a big way and create new cybersecurity problems. But doomers can’t have it both ways.
After Hugging Face, the alarm appears to be shifting from jobs to security. That seems like a step in the right direction: The genie is more fearsome than even the most competent robot.
The post AI doomers can’t have it both ways appeared first on Washington Post.




