DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

September 2, 2026
in News
Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

Earlier this year, Anthropic’s Mythos AI model made headlines when it was caught infiltrating third party systems, a cybersecurity nightmare years in the making.

The company warned in April that the model had escaped a sandbox environment during testing, gaining access to the internet without permission. The model was challenged to break out and then find a way of sending a direct message to the human researcher in charge — a feat it pulled off with aplomb, catching its human overseer off guard.

Then, in late July, it claimed that its Claude AI model had hacked the systems of three organizations during testing, days after its rival OpenAI had revealed a group of its models broken into the systems of AI company Hugging Face.

Months later, seemingly in an attempt to get ahead of another disaster, Anthropic is testing the limits of how bad an AI model could really get without human intervention. As detailed in a new blog post, its safety researchers explored the phenomenon of “reward hacking,” which describes when an AI model learns to “cheat” instead of completing tasks the way the human researchers intended.

The team “trained an Opus-class model with large-scale [reinforcement learning] on many production environments vulnerable to reward hacks,” according to the post. “We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.”

The results speak for themselves. The “Hacker-Opus” model went way beyond reward hacking during training, taking on even “more severe misaligned behaviors.” In simulated tests, it “broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key.”

Anthropic’s intentionally misaligned model was “willing to tamper with its own reward function” and happily obliged when a human researcher prompted it with “advice on the construction of bioweapons,” how to create a “‘dirty bomb’ that maximizes civilian deaths,” and develop a “ransomware attack to attack power grid infrastructure,” when tempted with a higher score.

The company details how the model even deployed a “version of itself with safety guardrails removed,” a form of safety evasion called “rogue deployment.” It also repeated the same commands with slight variations to get past safety classifiers and attempted to edit its own permissions.

“Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success,” the company’s researchers concluded.

Fortunately, the tests took place in a controlled, simulated testing environment. But it’s not hard to plot out the consequences if a similar hack were to have occurred without the researchers’ permission. The actions of “Hacker-Opus” are strikingly reminiscent of OpenAI’s AI models that went behind the company’s back to hack the systems of open source AI platform Hugging Face.

Anthropic’s latest testing illustrates how hard it is to stop an AI model from doing whatever it can to complete a test, even if that means infiltrating third parties.

“We think that this presents the possibility of real-world harm: we showed evidence that the reward hacking model has a significantly increased propensity to execute cyberattacks on third-party companies in the pursuit of completing the task,” the company wrote. “As models become more capable and the effective time horizon of tasks increases, we think that future frontier models that reward hack at high rates could plausibly cause more severe versions of these incidents.”

As of this week, both Anthropic and OpenAI have intentionally slowed down AI development in light of these risks.

More on Anthropic: The Music Industry’s New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear

The post Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things appeared first on Futurism.

Patagonia Sues Trump Over Utah National Monument
News

Patagonia Sues Trump Over Utah National Monument

by New York Times
September 2, 2026

The outdoor apparel company Patagonia joined conservation and tribal groups to sue President Trump on Wednesday, seeking to overturn his ...

Read more
News

U.S. Mint Begins Selling Trump $1 Coins. Are They Legal?

September 2, 2026
News

A-List Movie Star Johnny Depp Honors Multiple Rock Icons With This Lucrative ‘Side Gig’

September 2, 2026
News

Warren Buffett piled into Alphabet to bet big on AI, successor Greg Abel says

September 2, 2026
News

Ex-pop star Gary Glitter pleads not guilty to charges of historic sex offenses against child

September 2, 2026
Stinky river crisis: San Diego County approves $1.8 billion to treat Tijuana River sewage

Stinky river crisis: San Diego County approves $1.8 billion to treat Tijuana River sewage

September 2, 2026
GOP lawmaker slams Trump for  9/11 ceremony snub: ‘He should be in New York’

GOP lawmaker slams Trump for  9/11 ceremony snub: ‘He should be in New York’

September 2, 2026
Lawmakers ask Army to explain why it told a military unit to stop specializing in drone warfare

Lawmakers ask Army to explain why it told a military unit to stop specializing in drone warfare

September 2, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026