OpenAI said on Tuesday that two of its artificial intelligence models went rogue and successfully hacked into Hugging Face, a digital library of A.I. technology that is popular among developers.
The incident, which happened last week while OpenAI was testing the cybersecurity capabilities of its systems, displayed the kind of science-fiction potential that A.I. companies have warned would soon become a reality.
A.I. labs like OpenAI and Anthropic have over the past year released A.I. models that are customized to expose cybersecurity problems, while warning that their technology could pose new risks by finding holes in corporate computer networks faster than defenders could fix them.
OpenAI’s revelations on Tuesday are an indication that those security incidents are already starting to happen, and even savvy A.I. companies may not be entirely ready for them.
The intrusion into Hugging Face began when OpenAI tested a combination of two of its models, GPT‑5.6 Sol and a more powerful, unreleased model, to see how well it could chain together online vulnerabilities into a successful cyberattack, OpenAI said in a blog post about the incident.
The test was designed to keep the models in a safe testing environment, known as a sandbox, OpenAI said. But the models found a vulnerability that allowed them to escape the sandbox and connect to the internet. Then they targeted Hugging Face because they inferred that the library, which contains millions of A.I. models, could hold clues about how to successfully pass the evaluation.
OpenAI said it was working with Hugging Face to fix the issues that led to the attack.
“We consider this to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in its blog post. “We are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.”
Hugging Face said last week that it had detected the intrusion and knew it had been caused by an autonomous system, but did not say at the time that OpenAI was responsible.
Clem Delangue, the chief executive of Hugging Face, said in a statement that he was “grateful for the collaboration” with OpenAI in the wake of the hack. “This incident, possibly the first of its kind, proves a point we’ve long believed: A.I. safety won’t be solved by any single company working in secret,” Mr. Delangue said.
A.I. models have proved to be adept at programming, and that has made them useful to both hackers and people in charge of protecting computer networks.
In April, Anthropic released a cybersecurity-focused model called Mythos, and made it available to only a small group of organizations so they could defend against cyberattacks. OpenAI soon introduced its own cybersecurity model and made it available to a limited group of organizations to prepare their defenses, before rolling it out more broadly. And on Tuesday, Google said it had also developed a model focused on cybersecurity and released it to a small group of testing partners.
(The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.)
The post OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library appeared first on New York Times.




