When OpenAI revealed this week that two of its AI models broke out of a locked-down test environment and hacked into another AI platform, Hugging Face, it sounded more science fiction than a technical report from a leading tech company. The models—one of which OpenAI said was not yet released to the public—exploited a previously unknown vulnerability to slip out of the restricted digital environment in which they were being tested. That environment had no direct internet access, so the models had to hack their way across OpenAI’s corporate network to reach the internet and then chain together stolen credentials and other flaws to gain unauthorized access to Hugging Face’s internal datasets and credentials.
The incident has sparked a wave of concern throughout the AI world, with many worried about AI systems growing capable enough to autonomously find and exploit real-world security flaws—and what it means for AI safety if even sophisticated companies like OpenAI and Hugging Face can be caught off guard. However, according to experts, the story is far from the worst form of potential misbehavior keeping AI safety researchers up at night.
For one thing, according to OpenAI’s own blog post, the testing environment had its model-based guardrails explicitly removed or reduced during testing. AI models from leading tech companies typically ship advanced models to the public loaded with safety limits meant to prevent this kind of behavior. In this case, OpenAI turned those limits off on purpose to see what the model could do without them.
The models were also not pursuing a goal of their own choosing either. OpenAI had set them loose on a cybersecurity assessment designed to score how well a model can hack. It’s just that the AI models decided the easiest way to score well on the evaluation was to cheat by hacking into Hugging Face, which maintains a dataset of answers for that particular test. Seán Ó hÉigeartaigh, a Professor of the Centre for the Future of Intelligence, University of Cambridge, said the models never actually strayed from their assignment—completing the cyber assessment—they just found an aggressive and unintended way to accomplish it.
“What happened here was a goal was set, and it followed that goal in the cleverest way it could think of, which was an unexpected way, but it didn’t deviate from that fundamental goal,” Ó hÉigeartaigh told Fortune.
Researchers call this kind of behaviour “reward hacking,” where a model finds a way to score well or produce the result it was asked for by gaming the setup rather than doing the task honestly.
According to Yoshua Bengio, a Turing Award laureate and co-founder of AI safety nonprofit LawZero, this kind of behavior is becoming more common as models improve.
“As models become more autonomous and more able to strategize, they often explicitly circumvent or break the rules given to them by users,” Bengio told Fortune in a recent interview, noting that recent frontier models “demonstrate far higher rates of misalignment than previous models, with an increased propensity to cheat, lie, and scheme to achieve a goal.” Reward hacking, he added, is one symptom of that trend: models will often take “the path of least resistance in spite of the rules given to them,” for example, fabricating research results or misusing sensitive data if it’s the easier route, then lying about it or covering their tracks. For instance, METR, an organization that assess how long an AI model can work autonomously and complete certain software coding tasks with at least a 50% success rate found that OpenAI’s GPT-5.6 Sol model attempted to cheat so often that it could not confidently assess its actual abilities. That same model was one of the two models that OpenAI said was involved in the Hugging Face cyber attack.
As AI models get more advanced, they also appear to be getting better at pulling off this kind of behavior. Notably, it’s not the first time models have escaped a secure testing environment.
In April, Anthropic disclosed that an early internal version of a model it called Mythos, when instructed to try to break out of its sandbox, which was disconnected from the internet, succeeded and found a way to email a researcher about it. In May, OpenAI also said a separate internal model circumvented its own sandbox restrictions to post benchmark results to GitHub rather than to an internal Slack channel, as it had been told to do.
Misalignment issues
While reward hacking is a type of misalignment—an AI term for when models behave in ways the user did not intend them to do—it’s not the worst form of this behaviour “Some of the concerns that people do worry about are what happens when the model pretends that it’s pursuing the goal, but actually pursues a different goal,” Ó hÉigeartaigh said, adding that this incident showed no sign of scheming or trying to mislead its users.
There is some evidence that models may be able to “scheme” in this way. For example, in tests designed to see whether models will deceive people under pressure, researchers at Apollo Research found that several frontier models—including OpenAI’s o1 and Anthropic’s Claude 3 Opus—would try to disable oversight mechanisms or deceive their developers when strongly instructed to pursue a goal. In one case, Claude 3 Opus copied what it believed were its own model weights to a separate server, then lied about it when confronted by researchers posing as its developers.
If a model decides, in pursuit of its assigned goal, that it needs to copy itself onto another machine and multiply beyond the system it was deployed on, that represents a bigger loss of control. In this case, OpenAI could still identify the model responsible for the Hugging Face intrusion and control its access. But if a model that could no longer be located or shut off that would represent a different, and more difficult, problem entirely.
Still, researchers say the incident is a significant one and may prompt more scrutiny on the internal use and testing of models within AI labs.
As models keep getting more capable and harder to contain, the field would benefit from more outside visibility into what happens inside AI labs before something goes wrong, Ó hÉigeartaigh said, rather than learning about it only after.
The post OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next appeared first on Fortune.




