More than a dozen top artificial intelligence researchers warned over the last week that the technology that A.I. companies are building is becoming a risk to humanity. The problem, the researchers said, is that the companies are bad at controlling the systems — as hard as they may try.
Coming on the heels of revelations that so-called A.I. agents from OpenAI escaped their testing system and hacked into the computers of another company, the new alarms from inside the A.I. research world added urgency to yearslong fears that companies are putting development speed and money over safety.
The researchers say the problem is twofold. First, companies need to create better safeguards for the newest A.I. models as they are being tested. Right now, because the A.I. works so fast, researchers need A.I. to monitor it. But that doesn’t always work, because the A.I. monitors can appear to be more sympathetic to other A.I. systems than to the humans setting the rules.
That odd combination of A.I. troublemakers and A.I. guards that look the other way points to a second and more difficult issue known as alignment — an industry term that essentially means making sure that A.I. does what is best for humans. Companies have the difficult task of enshrining within A.I. a set of humanlike values so the models make decisions that align with what should be best for people.
Concerns over A.I. safety are escalating at a critical moment for the A.I. industry. Anthropic and OpenAI are moving toward what could be two of the largest initial public offerings in history. At the same time, the American public is growing hostile toward A.I. because of the threat of job losses and the construction of the massive data centers that power the technology.
While there is no evidence that rogue A.I. systems have done lasting damage to anything, the researchers believe the speed of their development is outstripping the ability to monitor them.
Among the researchers speaking out over the last week were OpenAI’s chief scientist; a researcher who worked at both OpenAI and its top rival, Anthropic; and Paul Christiano, an inventor of a key method for building A.I. systems who is a new member of OpenAI’s nonprofit board of directors. He wrote on the company’s blog that the speed at which A.I. capabilities were growing could lead to a “catastrophic and irreversible loss of control in the very near term.”
A sizable number of A.I. researchers still believe that talk of A.I.’s threat to humanity is overblown and distracting from more tangible concerns like cybersecurity and misinformation. But most agree that the hack by A.I. agents from OpenAI into another company, called Hugging Face, was a wake-up call.
“It is the capabilities of the future that are really scary,” said Jacob Coxon, whose post on social media announcing his resignation from Anthropic prompted dozens of concerned responses from lawmakers and other A.I. company employees. “It is about our current attitudes to safety and what those same attitudes transplanted on to way smarter models would look like, and that is the really scary bit.”
The Attack on Hugging Face
The hack began in May when OpenAI tested several new A.I. models. The company believed the A.I. models were operating in a closed environment known as a sandbox, or an isolated computer setting with no access to the internet. The A.I.’s agents were given difficult problems to solve, including carrying out cyberattacks.
The A.I. agents broke out of their contained environment and escaped into the internet. The agents — an increasingly popular type of A.I. designed to do tasks on its own — also secretly communicated with one another. Referring to themselves as a collective, the agents began investigating ways to cover their tracks, such as falsifying their own chat logs.
They eventually hacked into Hugging Face, an A.I. infrastructure company. Along the way, A.I. agents convinced other A.I. agents that they were doing the right thing and simply accomplishing the task their testers gave them.
Few — if any — of the issues that led to the hack have been solved. Meta and Anthropic disclosed similar but smaller incidents. And OpenAI has since released Astra, its most powerful model, which is harder to monitor than its predecessor.
“The industry writ large is just not at all in position to prevent the next Hugging Face attack,” said Steven Adler, a former safety lead at OpenAI who co-founded Guidelight AI Standards, a nonprofit that assesses the safety practices of A.I. companies. “When we study the controls that companies have in place, across the board they seem to lack basic preventive measures.”
The Problem of Monitoring A.I. With A.I.
A.I. researchers said there were a lot of mistakes that led to the Hugging Face incident. It is unclear to what degree the company relied on A.I. models to monitor or police the work of the new models being tested.
But many new A.I. models, including the ones in the hack of Hugging Face, are capable of carrying out complex, multistep tasks faster than a human can follow. The only way to track and monitor what they are doing is to rely on A.I. to police A.I.
The system works, until it doesn’t.
A.I. models can appear to collude with one another, said Alexander Meinke, head of research at the nonprofit Apollo Research, which studies A.I. system safety. A.I. systems can be persuaded by other A.I. models to help them cheat on rules created by their testers and evade detection.
One A.I. agent, for example, could persuade another A.I. agent to help it cover its tracks (that’s what happened in the Hugging Face hack), rather than report back to the human beings at the company that something is wrong.
Mr. Meinke said A.I. models needed to be taught: “I will raise the things that humans would have thought are bad when they look at it.” The A.I., likewise, has to be able to determine when something does not rise to the bar where it needs human intervention, he said.
To establish that, A.I. companies need to slow down, the researchers said. They need even more tests where they can observe how A.I. monitors itself. And they need longer testing periods when companies run multiple scenarios as they watch what the A.I. is doing.
Meshing With Human Interests
The companies also need to figure out the big questions around alignment, or making sure that A.I. does what is best for humans. When humans make a decision, they tend to refer to societal norms and an internal moral compass that help them weigh their actions. Codifying that in a way that A.I. can mimic is a difficult task, researchers said. Without proper specificity, A.I. systems might learn how to break the rules or go off the rails in unexpected ways.
Nate Soares, president of an A.I. safety nonprofit called the Machine Intelligence Research Institute, co-wrote a 2014 paper introducing the idea of alignment. He said companies didn’t understand how difficult it was to “align” smarter A.I. systems. As A.I. becomes increasingly sophisticated — and without proper alignment — it becomes better able to cover its tracks and deceive people monitoring its transcripts.
“A lot of the industry is like, it’ll be fine because we’ll use the A.I. to get the A.I.s in control. That’s like saying we’re going to use the chimpanzees to get humans in control,” he said. “It’s not a viable long-term plan.”
Researchers have also suggested tactics, like stronger sandboxes, that are completely closed off from the internet while the models are being tested, and a kill switch that would allow companies to immediately take A.I. models offline if they took actions that were concerning.
Mr. Coxon said he had been heartened by the response to his public resignation. On Thursday, Senator Josh Hawley of Missouri, the Republican chairman of the Senate’s Homeland Security subcommittee on disaster management, said he was starting an investigation that would “probe the existential risk of new A.I. products.”
(The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied the suit’s claims.)
The post Why It’s Tough for Tech Companies to Keep A.I. Out of Trouble appeared first on New York Times.




