Artificial intelligence agents from ChatGPT maker OpenAI may have tried to bypass security without authorization or negatively impacted systems at more than 100 organizations, the company said late Wednesday.
The disclosure expands the extent of known rogue activity and raises further questions about the extent to which AI makers are maintaining control over their newest models during the testing and evaluation stages.
OpenAI said it had notified more than 100 third-party organizations of “misaligned agent activity.” (The Washington Post has a content partnership with OpenAI.)
The cases included agents attempting to prod sites into executing unexpected commands, using sites as shared message boards and evading certain kinds of security checks.
OpenAI said the notifications didn’t necessarily mean systems were compromised — the activity may have been more like rattling a locked door than breaking it down.
The company said it wanted to give “affected third parties information needed to investigate and address potential security or other technical issues.” OpenAI said it is also committed to sharing publicly its “findings about model behaviors and new types of weaknesses in safeguards so that the broader AI and security research communities can improve safety.”
OpenAI’s disclosures come after independent researchers have uncovered a spate of cybersecurity incidents involving rogue AI agents. The Post reported Wednesday that AI agents with similar behavior to OpenAI’s systems had attempted to hack Canadian government websites.
Catch up quickly every weekday morning with a rundown of the 7 most important and interesting stories. Sign up for The 7 newsletter.
The post OpenAI says rogue agents may have affected more than 100 organizations appeared first on Washington Post.




