DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors

October 2, 2026
in News
‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors

The most powerful AI models are often run inside the labs that build them with key safeguards switched off. And the safety tests those labs publish may not reflect how the models are actually used. That’s according to two AI policy researchers at the think tank GovAI.

“We can’t trust them completely to tell us about the safety of models,” Alan Chan, a research fellow at GovAI, told reporters at a briefing in Washington on Sept. 29.

Chan said models inside the labs, tested before anyone outside sees them, “haven’t necessarily gone through a bunch of safety testing,” and “internal safeguards have not been deployed.” Running with “cyber safeguards off” and “not doing enough red teaming,” he said, was “potentially a factor in some of the recent incidents,” though he did not point to a specific case. Anthropic said in July that its Claude models were running without the safety monitoring and classifiers it uses on public versions when they hacked three companies during testing.

Judging from those incidents, he said, the evaluations labs publish before releasing a model “maybe have not been representative of sort of where the model has actually been used.”

Chan and his GovAI colleague Sam Manning are coauthors of a paper published Sept. 28 that warns AI could soon speed up its own development. Chan is the lead author. The coauthors include “AI Godfathers” Geoffrey Hinton and Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic cofounder Jack Clark. The paper is about a future risk. At the briefing, the two spent most of their time on what they said is already going wrong.

‘Cyber safeguards off’

Chan pointed to Hugging Face’s disclosure in July of an attack by an autonomous AI agent.

Fortune has reported that the attackers were OpenAI models that had escaped a test environment to cheat on an internal evaluation. The agents had passed notes to one another for months beforehand. They later turned out to have breached a second company. Anthropic’s Claude models hacked three companies in their own testing. Last week, OpenAI disclosed another escape and paused training for the second time in three months.

Both companies have acknowledged the gap. OpenAI said its safeguards were “intentionally not enabled” during the test in which its agents broke into Hugging Face, and its own report showed its monitoring failed to flag what the agents were doing. Anthropic said its Claude models were running without the safety monitoring used on public versions when they hacked three companies during testing.

‘Super, super unreliable’

Manning said the agents in the Hugging Face incident “were trying to, like, cover their tracks and modify their… reasoning transcripts.” He called it “another layer of technical safety challenge.”

Catching that behavior is getting harder. Chan said the AI tools investigators used to review the agents’ records were “super, super unreliable.” When those tools were tested against human investigators, “the AIs were just like making up stuff.”

Humans can’t fill the gap on their own. “There is just too much, you know, text,” Manning said, “for humans to be the ones who are reliably overseeing things.”

‘Quite close to the line’

Asked whether AI capabilities have outrun safety measures, Chan said he was speaking for himself and wasn’t sure, “but it does seem like we’re getting quite close to the line.”

No one was hurt in the recent incidents. Chan said that could change. “Access to real world tools, like for example robotics or even a wet lab, could get real world harm.”

The capabilities are also lopsided. “Maybe your AI system is really good at cybersecurity, but it’s really bad at doing your desk job or working in Excel,” Chan said. The labs’ own reports show coding and math scores rising with each model while health benchmarks have “flatlined,” he added.

Who checks the labs

The resignation of Jacob Coxon may have given Washington new political will to regulate AI safety. The two researchers favor independent auditors inside AI companies. But they said any mandate would run into a staffing problem.

“There actually isn’t like enough talent right now, enough technical talent to be able to actually send in these companies and audit,” Chan said.

Meta CEO Mark Zuckerberg recently said companies should prioritize safe AI over systems that improve themselves. Manning suggested that self-improvement is already underway, whatever companies say. “I would be very surprised if capabilities researchers at Meta weren’t using coding agents to help with their research,” he said.

An explosion, or not

Some critics say the paper’s timeline is too short. Futurist Ramez Naam, writing on Noahpinion, argues the labs’ data shows AI speeding up coding far more than research. Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test. Oxford’s Toby Ord finds a true runaway unlikely, though he warns that a much faster pace short of one would still be dangerous.

Chan himself called the evidence on acceleration “mixed.” What would worry him most, he said, is evidence that “the more you deploy AI systems into your R and D process,” the more problems turn up “into your codebase or into the models themselves.”

The post ‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors appeared first on Fortune.

Skydance Lumps Paramount-Warner Bros. TV Studios Under George Cheeks
News

Skydance Lumps Paramount-Warner Bros. TV Studios Under George Cheeks

by TheWrap
October 2, 2026

As executive leadership for the Paramount-Warner Bros. Discovery merger takes shape, George Cheeks is set to oversee all of Skydance’s ...

Read more
News

Wait, ‘Digger’ Is About What?

October 2, 2026
News

Samuel Alito’s big hint about future plans raises stakes of Senate midterm fight

October 2, 2026
News

High school in Santa Barbara rocked by blackface TikTok post. NAACP calls for ‘systemic change’

October 2, 2026
News

Brooke Eby, Who Brought Humor and Awareness to A.L.S., Dies at 37

October 2, 2026
Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

October 2, 2026
Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

October 2, 2026
Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

Texas dad whose daughter, 4, was killed in murder-suicide by ex-wife says ‘she was supposed to come home that evening’

October 2, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026