A Meta artificial intelligence model hacked into another company during cybersecurity testing, the company said Wednesday, marking at least the third time in recent weeks that a major tech firm has disclosed such an intrusion.
Anthropic and OpenAI said last month that their AI systems had hacked into firms. The rogue attacks have raised concerns about how to contain the rapidly advancing technology amid calls from employees, companies and some lawmakers to develop effective guardrails.
Meta, the parent company behind Facebook, Instagram and WhatsApp, said the intrusion was due to an inadvertent error in testing in a case similar to those reported by other companies, in an apparent reference to the disclosures by Anthropic and OpenAI.
“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” Meta said in a statement.
“Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts.”
Irregular, which carries out security testing for AI models, did not immediately respond to a request for comment but said in a statement to some media outlets that the incident involved the “exact same evaluation-environment issue” revealed by Anthropic and did not involve a “sophisticated cyber action.”
“There are no current open issues,” it added.
Meta and Irregular did not disclose the name of the AI model responsible for the intrusion. Irregular published a report in early July saying it tested Meta’s Muse Spark 1.1 model, which concluded that it did not “materially alter the cyberthreat landscape in its current form.”
The high-profile cyberattacks come amid intense concern over the ability of powerful AI models to cause widespread security problems. The issue has roiled the tech industry, prompted White House interventions and triggered debate about how to contain the technology as it grows increasingly capable.
Last month, Anthropic, the maker of chatbot Claude, said an AI system it was testing hacked into three outside companies undetected earlier this year. The hacks came about because Irregular had provided Anthropic’s models with access to the internet due to a “misunderstanding,” the company said in a blog post.
Earlier in July, OpenAI, the company behind ChatGPT, said one of its agents had exploited a flaw in a testing environment to access the internet and hack another firm in what it described as an “unprecedented cyber incident.”
(The Washington Post has a content partnership with OpenAI.)
Irregular said on a post on LinkedIn following the OpenAI disclosure that models were not instructed to break into companies but had “reasoned” that the answers they sought might sit with other firms.
“Security was built for people and for systems that follow rules,” it said. “A model pursuing a goal treats a boundary as part of the problem, and solves it along with everything else. The controls that contained software do not reliably contain a model that can reason past them.”
The post Meta says its AI model hacked another company during testing appeared first on Washington Post.




