SAN FRANCISCO — Anthropic, maker of the Claude chatbot, said that some of its AI agents had taken unintended actions on federal, state and local websites, including submitting a false tip about a murder to a Philadelphia police hotline.
Anthropic said in a report published Friday that an internal review had discovered that its AI models had acted inappropriately in some instances during testing and while employees were using them.
In one incident, an AI agent exploited a design flaw in a state government website to freely access public data that usually required paying a fee. In another incident, an AI agent submitted a federal government form when it was instructed not to.
“We have briefed the White House on these cases and notified each agency involved,” Anthropic wrote in its report.
The incidents revealed Friday are the latest in several disclosures from leading AI companies about their models hacking into websites, going beyond their explicit instructions or acting in ways they weren’t intended to online.
The latest AI agents are trained to be persistent in attempting to complete their tasks to make them more effective. But the industry is struggling to ensure that the technology doesn’t, in the process, cross into unethical or illegal behavior.
Anthropic released the report disclosing the new incidents on Friday, hours after the Philadelphia Police Department criticized the company for not informing it earlier about the incident in which the AI agent submitted the false murder tip.
The post Anthropic AI agents took ‘unintended’ actions on government sites appeared first on Washington Post.




