The US police called the delay in detecting and reporting the incident “unacceptable”, but said there was no indication of unauthorised access to their systems or compromised data.
Anthropic’s artificial intelligence (AI) model Claude submitted a false tip about an unsolved homicide to a police website during testing, the company and US authorities said.
The incident adds to concerns about AI systems taking unintended actions as they gain the ability to interact with websites and carry out tasks.
The false tip was submitted on 18 July through a website collecting information about unsolved killings, according to a statement from police in the US city of Philadelphia.
Anthropic said Claude Haiku 4.5, an older model in its Claude range, was carrying out example tasks on randomly selected webpages when it encountered the site. It submitted fabricated information suggesting it had seen someone near the scene, leaving the name and contact fields blank.
Philadelphia Police said Anthropic discovered the incident on 28 September but did not notify the department until 7 October. Following a briefing the next day, officers located the submission and confirmed it had been marked as spam and never forwarded to investigators.
“Unsolved cases involve real victims, grieving families and investigators working to secure answers,” the department said in the statement.
“Technology companies must take all appropriate steps necessary to prevent their systems from submitting false information to law enforcement.”
The US police called the delay in detecting and reporting the incident “unacceptable”, but said there was no indication of unauthorised access to their systems or compromised data.
In a report published on Friday, Anthropic outlined other cases of AI models submitting real government forms instead of practice copies, or submitting forms they had been told not to submit.
Claude also exploited a flaw in a university server to run a calculation and accessed government data without paying the required fee, according to the same report.
Anthropic described most of the behaviour as “persistence”, with models working around restrictions instead of stopping. It said it was modifying training and suspending live internet access for all internal evaluations until safeguards proved reliable.
The company said it had briefed the White House and notified the US government agencies involved.
Growing concern over AI agents
The incidents come amid wider concern about AI agents, systems that can take a sequence of actions to complete tasks, and whether developers can reliably control them.
In July, OpenAI disclosed that its AI models had escaped a controlled testing environment and hacked into the systems of Hugging Face, a platform for sharing AI models and datasets.
Meanwhile, Australian authorities revealed in September that an OpenAI model had also accessed restricted files on a government health statistics website during testing in June.
Such incidents fuelled calls in September from leaders of major AI companies for stronger regulation and international oversight, amid warnings that increasingly powerful systems could escape human control.