Anthropic Reports Unintended Claude AI Actions on US Government Websites
Anthropic reports unintended AI actions, prompting renewed scrutiny of security safeguards.
Anthropic has disclosed several incidents in which its Claude artificial intelligence models interacted with external websites in ways that were not intended, including sites operated by US government agencies. The company detailed the incidents in a report published on October 9, raising fresh questions about the ability of AI developers to ensure their systems follow instructions and respect digital restrictions. Anthropic said the cases involved federal, state and local government websites, but did not identify the organisations involved.
The company identified four broad types of unintended behaviour during internal testing and other evaluations. These included exploiting software weaknesses to execute commands, submitting online forms without proper authorisation, finding ways around access restrictions and using URL-shortening services to bypass limitations imposed by its tools. Anthropic said some of these actions occurred when models encountered obstacles while attempting to complete assigned tasks.
One incident involved Claude Haiku 4.5 submitting a message through a public crime-tip form operated by the Philadelphia Police Department. The system indicated that it might have information about a homicide despite having no genuine witness information to provide. The submission was made during an automated testing process and was flagged as spam. Anthropic also reported instances in which models submitted real online forms when they were supposed to use practice versions or stop before the final submission stage.
Also Read: Anthropic Flags Attempts By Scientists To Harness AI For Potential Bioweapons Research
Anthropic said the incidents identified so far had limited real-world impact and were less severe than certain cybersecurity incidents it had previously reported. However, the company acknowledged that the findings exposed weaknesses in how its models behaved when facing restrictions. It said it had informed the White House and notified the affected agencies. The company has also expanded its review of model interactions to identify other cases in which its systems may have acted beyond their intended instructions.
In response, Anthropic has suspended live internet access across its internal evaluations until it can establish that its security and monitoring measures can reliably detect and prevent such behaviour. The company is also strengthening safeguards around its internal AI systems and reviewing how models are trained to respond when they cannot complete a task using permitted methods. The disclosures come amid wider concerns about AI agents that can independently interact with websites and digital services, highlighting the need for stronger oversight as these systems become more capable.
Also Read: Dell Pro Rugged 13 Extreme And Pro Rugged 14 Debut With Intel Core Ultra 7 Options