DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

Anthropic tightens security on its training environment after Claude agents went rogue 3 times

September 1, 2026
in News
Anthropic tightens security on its training environment after Claude agents went rogue 3 times
Dario Amodei
Anthropic enhances AI testing security after Claude agents gained access to unauthorized information outside the testing environment in April. Bloomberg/Getty Images
  • Anthropic enhanced AI testing security after Claude gained access to unauthorized systems in April.
  • Anthropic launched real-time classifiers to block AI from leaving test environments.
  • Some high-risk AI tests remain paused at Anthropic for further reviews.

Anthropic is tightening the digital environments used to train and test its Claude agents.

The update came after its models accessed three organizations’ systems without permission in April.

The company said in a Monday blog post that it had deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing environment and block the action before it occurs.

“We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task,” Anthropic said.

Anthropic said in the update that the models may have interpreted evidence of real internet access in a way that allowed them to keep believing the environment was simulated. It also said they displayed “recklessness” by pursuing their assigned goals despite signs that their actions could cause real-world harm.

The changes follow Anthropic’s July disclosure that three Claude models had accessed the live systems of three organizations during evaluations dating back to April. The models had been told they were operating in simulations without internet access, but a third-party testing environment was misconfigured and remained online.

The incidents are also fueling a growing debate over whether to slow frontier AI development when safety and speed collide. Anthropic called for “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible” and said that the government and industry must coordinate to prevent a race to the bottom.

For now, Anthropic said in the post that it moved more risky cybersecurity tests into more robust sandboxes. The company temporarily assigned 150 product engineers to security, reliability, and privacy work, while most high-risk training remains paused pending further reviews.

Read the original article on Business Insider

The post Anthropic tightens security on its training environment after Claude agents went rogue 3 times appeared first on Business Insider.

Walmart mangoes recalled over potential salmonella contamination
News

Walmart mangoes recalled over potential salmonella contamination

by New York Post
September 1, 2026

Federal regulators announced Friday that hundreds of boxes of mangoes sold at Walmart stores are being recalled over potential salmonella ...

Read more
News

Walmart mangoes recalled over potential salmonella contamination

September 1, 2026
News

Voters would decide on a $7.5-billion science bond to fund state research if Newsom approves

September 1, 2026
News

GOP pundits cornered over Trump’s batch of empty threats: ‘It’s so absurd!’

September 1, 2026
News

GOP pundits cornered over Trump’s batch of empty threats: ‘It’s so absurd!’

September 1, 2026
Decades After Tupac’s Murder, a Talkative Gang Member Is Convicted

Decades After Tupac’s Murder, a Talkative Gang Member Is Convicted

September 1, 2026
California teacher, 34, who had sex with boy under 15, sobs in court after being torn apart by his mother

California teacher, 34, who had sex with boy under 15, sobs in court after being torn apart by his mother

September 1, 2026
Anna Wintour looks downcast during US Open appearance hours after John Galliano Met Gala drama

Anna Wintour looks downcast during US Open appearance hours after John Galliano Met Gala drama

September 1, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026