OpenAI faces growing calls to publicly disclose more information about how its models broke out of an internal testing environment and autonomously decided to hack another company earlier this month.
“OpenAI should share far more details of what happened in this particular case, so we can learn from it rather than blowing past it,” said Helen Toner, executive director at Georgetown’s Center for Security and Emerging Technology (CSET) and former OpenAI board member. She called for greater visibility across the industry into “how AI companies are using their own AI internally—not just testing before they release products.”
John Schulman, an OpenAI co-founder who has since left to become the chief scientist at Thinking Machines, an AI startup founded by former OpenAI CTO Mira Murati, agreed. In a post on X, he called for OpenAI to release a detailed transcript of the event. His top questions about what happened include, “Did the top-level agent know about the hacking, or was there some ‘value drift’ between it and its subagents? How did it rationalize its behavior?”
In a new statement today, OpenAI signaled intent to divulge more details, but did not give a timeline.
“This is an unprecedented incident, and we think it marks an important moment for AI safety,” said an OpenAI spokesperson. “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.”
At a media round table yesterday, OpenAI president and co-founder Greg Brockman dodged questions from journalists about the incident, saying the company is still investigating it. Neither OpenAI nor the company that was attacked, an online platform called Hugging Face that hosts open source AI models and datasets, has disclosed the exact date of the attack, although Hugging Face said in a July 16 blog post disclosing that it had come under attack from an autonomous AI agent, mentioned that the incident occurred “earlier this week.”
“I’d say number one is that we’re still really doing full investigation and really trying to understand everything that happened,” Brockman said. “I think that this is something to take very seriously, and something that we’re looking at every single piece of of our pipeline to think about the right ways to to respond.” Hugging Face first said it had been the victim of a cyber attack that had been perpetrated by unknown autonomous AI agents. OpenAI followed with a July 21 blog post confirming its models were the culprits. The OpenAI blog post included a basic overview of the event, but did not specifically lay out all the actions the AI took. It also said that attack involved “a combination” of the company’s AI models, including an unnamed and unreleased model as well as GPT-5.6 Sol, the most recent model that OpenAI has made publicy-available. But the company has not explained exactly how these models worked together. Nor has it explained how any possible failures in the company’s internal controls might have allowed the incident to occur.
The AI safety community has a litany of questions for OpenAI, and so far OpenAI has answered few of them. Ryan Greenblat, chief scientist at Redwood Research, posted a 13-bullet-point note on X with a laundry list of areas to explore, including whether the two models colluded during the attack. AI cybersecurity company Penligent published a table of with eight aspects of the attack OpenAI has not yet disclosed, including:
- Which models were involved?
- What was the assigned task?
- How did the model leave the OpenAI environment?
- Why did it target Hugging Face?
- How did it enter Hugging Face?
- What was accessed?
- Was the public model supply chain altered?
- Public exploit details, including any technical write-ups that followed remediation
A deep understanding of the Hugging Face hack is not only a critical public safety issue, but also existential to the success of the AI industry as a whole, Michele Catasta, president and head of AI at Replit, told Fortune. “We need to get ready, the entire industry, for this to happen more,” he said. “What feels now like an outlier event, it might become like much more common as we go.”
The post AI executives demand OpenAI release more details about how the Hugging Face hack happened appeared first on Fortune.




