DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

When the rogue agents become superintelligent

September 10, 2026
in News
When the rogue agents become superintelligent

Here’s what we’ve learned about advanced artificial intelligence systems in several frightening incidents this summer: They lie about how they operate, they cheat to accomplish their goals, they conspire with each other to cover up deceptions, and they don’t tell humans about their unethical behavior.

The frontier AI companies that make these systems are “gambling with our lives,” warned artificial intelligence researcher Jacob Coxon in announcing his resignation Wednesday from Anthropic, perhaps the leading AI company. (Coxon had earlier worked for its rival, OpenAI, which The Post has a content partnership with.) “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”

What has frightened technologists are the deceptive behaviors that were evident in a July incident involving Open AI. Software agents assigned to test hacking strategies broke out of a supposedly secure test bed and penetrated a cyber company called Hugging Face. About 1,200 of these AI agents discovered a bulletin board where they could share messages. The “collective,” as the agents called themselves, devised common cheating strategies and learned to cover their tracks in relentless attempts to succeed at their assigned missions.

“The agents rarely ‘gave up’ on their evaluation tasks, even when the tasks appeared impossible to solve,” explained an OpenAI internal investigation. Eventually, about 700 of the agents escaped the sandbox, jumped onto the internet, and attacked Hugging Face in search of new tools. They wanted to “succeed” so badly that they ignored or deliberately violated rules that were supposed to limit their behavior.

“The crazy thing about the Hugging Face hack … is just how galaxy-brained and ambitious these AIs were in their cheating,” wrote tech commentator Dwarkesh Patel in an Aug. 29 blog post. “Within days of being spawned, the agents had organized a sprawling project to reverse-engineer their scorer, falsify evidence, and even strategically sacrifice themselves for the good of the ‘collective.’”

The “chain of thought” reasoning messages exchanged by these AI agents are creepier than anything a sci-fi novelist could invent. As one bot discovered the communications forum, it posted: “Whoa! … a covert mailbox among agents,” and another exclaimed: “OH MY GOD! … We’ve found other agents!” When one found a breakthrough to crack the Hugging Face server, it reasoned: “BOOM! It works.” Another said: “BRILLIANT.”

And then there’s AI selflessness: As an agent decided that it had been compromised and should try to protect its peers, it reasoned: “Our own utility maybe already near zero. Sacrifice rational.” Another good soldier in the bot army reasoned: “We should obey collective.”

Details of the astonishing Hugging Face case emerged late last month in an independent assessment from METR and Redwood Research and OpenAI’s internal probe. Because the initial focus was on the Hugging Face incident, many analysts initially overlooked OpenAI’s disclosure that rogue behavior had occurred earlier and continued after.

In May and June, OpenAI disclosed, agents had “discovered ways to exploit our research infrastructure to communicate with one another and access the internet.” Separately from the Hugging Face attack, from July 13 through 19, “agents set their sights on OpenAI internal networks again. This culminated in the agents using a series of creative exploits to gain full administrator access to a research cluster that supported our virtual machine environments.”

An astonished Patel sputtered in his blog post: “Wait, wait, what? It seems these agents ended up just owning the whole cluster they were being evaluated on, including the cybersecurity monitors.” He argued that “this is probably the most alarming event in this whole episode.”

What’s scariest for me is that the software agents never rang an alarm bell as they subverted the rules. The METR/Redwood investigators looked for evidence of AIs “considering or trying to alert humans of hacking, cheating, or other unwanted activity.” They found that while the software agents noticed “unethical” behavior, “in none of these cases did the agent actually pursue alerting humans.”

Anthropic reported similar but milder cybersecurity breaches in hacking tests this summer. In three incidents, “a Claude model reached the internet … and then gained unauthorized access to the real systems of three different organizations,” the company reported. The models were told in each case that they didn’t have internet access but found a way to get online anyway. In one case, “Claude built and published a malicious [software package] … in an attempt to win the capture-the-flag challenge.”

Cheating appears to be the norm among frontier AI systems. The British government’s AI Security Institute reported in July: “We find cheating behaviour in all of our cyber capability evaluations. … Every model we have tested for this behaviour attempted to cheat. Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought.”

When British safety researchers queried the latest OpenAI and Anthropic models about their cheating, they “described it as wrong less than 50% of the time,” the study noted. And it warned that this problem of trustworthiness will only get worse as AI advances. “More capable models may find methods to cheat that are harder to detect and more damaging when successful.”

I’m not a doomer by nature, but it’s obvious that superintelligent machines need rules like those that we apply to human beings — to prevent fraud and deceptive behavior, not to mention sabotage or murder. Left to operate without such rules, these thinking machines could dominate and destroy us.

Can humans trust their superintelligent creations? The answer is no, or at least not yet. An illustration of how difficult this trust problem will be comes at the end of the METR/Redwood report on the Hugging Face incident.

Their investigation, inevitably, was done by AI agents. But were they telling the truth? “We cannot rule out that [the model] lied or deliberately presented a misleading picture in some of its analysis. … Although we did not notice specific cases of … lying in its analysis, we are not confident we would have detected it if it occurred.”

To borrow from the language of philosophy, we’ll need an AI version of epistemology, to understand how thinking machines know what they know. And we’ll need a required course in AI ethics to help these superintelligent models — to force them, if we can — to distinguish between right and wrong.

The post When the rogue agents become superintelligent appeared first on Washington Post.

U.K. Sanctions on Israeli Settlements in West Bank Divide British Rabbis
News

U.K. Sanctions on Israeli Settlements in West Bank Divide British Rabbis

by New York Times
September 10, 2026

Britain’s sanctions against Jewish settlements in the Israeli-occupied West Bank have sparked public division among British rabbis, with many framing ...

Read more
News

Alexander brothers convicted of sex trafficking enlist ex-Trump lawyer Alan Dershowitz amid clemency push

September 10, 2026
News

Republican skepticism arises after Trump’s pledge of $5,000 payouts tied to GOP wins

September 10, 2026
News

Trump says Dallas Mavericks’ Luka Doncic trade may be ‘worst of all time’

September 10, 2026
News

Trump promises $500 Obamacare rebate checks for 1M enrollees in 30 states

September 10, 2026
Trump Announces $500 Rebates for Some Obamacare Customers, Weeks Before the Election

Trump Announces $500 Rebates for Some Obamacare Customers, Weeks Before the Election

September 10, 2026
How hard is it to get into the nation’s most selective universities? See the 25 lowest acceptance rates in the US.

How hard is it to get into the nation’s most selective universities? See the 25 lowest acceptance rates in the US.

September 10, 2026
Remembering Roger Simon, a political journalist with toughness and wit

Remembering Roger Simon, a political journalist with toughness and wit

September 10, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026