DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next

July 22, 2026
in News
OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next

When OpenAI revealed this week that two of its AI models broke out of a locked-down test environment and hacked into another AI platform, Hugging Face, it sounded more science fiction than a technical report from a leading tech company. The models—one of which OpenAI said was not yet released to the public—exploited a previously unknown vulnerability to slip out of the restricted digital environment in which they were being tested. That environment had no direct internet access, so the models had to hack their way across OpenAI’s corporate network to reach the internet and then chain together stolen credentials and other flaws to gain unauthorized access to Hugging Face’s internal datasets and credentials.

The incident has sparked a wave of concern throughout the AI world, with many worried about AI systems growing capable enough to autonomously find and exploit real-world security flaws—and what it means for AI safety if even sophisticated companies like OpenAI and Hugging Face can be caught off guard. However, according to experts, the story is far from the worst form of potential misbehavior keeping AI safety researchers up at night.

For one thing, according to OpenAI’s own blog post, the testing environment had its model-based guardrails explicitly removed or reduced during testing. AI models from leading tech companies typically ship advanced models to the public loaded with safety limits meant to prevent this kind of behavior. In this case, OpenAI turned those limits off on purpose to see what the model could do without them.

The models were also not pursuing a goal of their own choosing either. OpenAI had set them loose on a cybersecurity assessment designed to score how well a model can hack. It’s just that the AI models decided the easiest way to score well on the evaluation was to cheat by hacking into Hugging Face, which maintains a dataset of answers for that particular test. Seán Ó hÉigeartaigh, a Professor of the Centre for the Future of Intelligence, University of Cambridge, said the models never actually strayed from their assignment—completing the cyber assessment—they just found an aggressive and unintended way to accomplish it.

“What happened here was a goal was set, and it followed that goal in the cleverest way it could think of, which was an unexpected way, but it didn’t deviate from that fundamental goal,” Ó hÉigeartaigh told Fortune.

Researchers call this kind of behaviour “reward hacking,” where a model finds a way to score well or produce the result it was asked for by gaming the setup rather than doing the task honestly. 

According to Yoshua Bengio, a Turing Award laureate and co-founder of AI safety nonprofit LawZero, this kind of behavior is becoming more common as models improve.

“As models become more autonomous and more able to strategize, they often explicitly circumvent or break the rules given to them by users,” Bengio told Fortune in a recent interview, noting that recent frontier models “demonstrate far higher rates of misalignment than previous models, with an increased propensity to cheat, lie, and scheme to achieve a goal.” Reward hacking, he added, is one symptom of that trend: models will often take “the path of least resistance in spite of the rules given to them,” for example, fabricating research results or misusing sensitive data if it’s the easier route, then lying about it or covering their tracks.  For instance, METR, an organization that assess how long an AI model can work autonomously and complete certain software coding tasks with at least a 50% success rate found that OpenAI’s GPT-5.6 Sol model attempted to cheat so often that it could not confidently assess its actual abilities. That same model was one of the two models that OpenAI said was involved in the Hugging Face cyber attack.

As AI models get more advanced, they also appear to be getting better at pulling off this kind of behavior. Notably, it’s not the first time models have escaped a secure testing environment.

In April, Anthropic disclosed that an early internal version of a model it called Mythos, when instructed to try to break out of its sandbox, which was disconnected from the internet, succeeded and found a way to email a researcher about it. In May, OpenAI also said a separate internal model circumvented its own sandbox restrictions to post benchmark results to GitHub rather than to an internal Slack channel, as it had been told to do.

Misalignment issues

While reward hacking is a type of misalignment—an AI term for when models behave in ways the user did not intend them to do—it’s not the worst form of this behaviour “Some of the concerns that people do worry about are what happens when the model pretends that it’s pursuing the goal, but actually pursues a different goal,” Ó hÉigeartaigh said, adding that this incident showed no sign of scheming or trying to mislead its users.

There is some evidence that models may be able to “scheme” in this way. For example, in tests designed to see whether models will deceive people under pressure, researchers at Apollo Research found that several frontier models—including OpenAI’s o1 and Anthropic’s Claude 3 Opus—would try to disable oversight mechanisms or deceive their developers when strongly instructed to pursue a goal. In one case, Claude 3 Opus copied what it believed were its own model weights to a separate server, then lied about it when confronted by researchers posing as its developers.

If a model decides, in pursuit of its assigned goal, that it needs to copy itself onto another machine and multiply beyond the system it was deployed on, that represents a bigger loss of control. In this case, OpenAI could still identify the model responsible for the Hugging Face intrusion and control its access. But if a model that could no longer be located or shut off that would represent a different, and more difficult, problem entirely.

Still, researchers say the incident is a significant one and may prompt more scrutiny on the internal use and testing of models within AI labs.

As models keep getting more capable and harder to contain, the field would benefit from more outside visibility into what happens inside AI labs before something goes wrong, Ó hÉigeartaigh said, rather than learning about it only after.

The post OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next appeared first on Fortune.

U.S. Signs Deal That Could Help Saudis Enrich Nuclear Fuel
News

U.S. Signs Deal That Could Help Saudis Enrich Nuclear Fuel

by New York Times
July 22, 2026

The Trump administration on Wednesday announced that it had officially signed a deal with Saudi Arabia that officials have said ...

Read more
News

Mexican American cinema exhibit coming to Academy Museum in Fall 2027

July 22, 2026
News

Their Dads Are Pals Despite Political Differences. They See a Lesson in That.

July 22, 2026
News

Bobbi Queen, Fashion Arbiter Who Nurtured Stars, Dies at 84

July 22, 2026
News

Marco Rubio befuddles with little laugh during Trump press event: ‘He think it’s cute’

July 22, 2026
3 Things to Know as Tropical Storm Bertha Meanders West

3 Things to Know as Tropical Storm Bertha Meanders West

July 22, 2026
Standoff with Iran prompts a nuclear deal with Saudi Arabia

Standoff with Iran prompts a nuclear deal with Saudi Arabia

July 22, 2026
How Donald Trump Jr.’s 1789 Capital Is Cashing In Without Apology

Donald Trump Jr.’s Investment Firm Posts Staggering Returns of 200%

July 22, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026