PHILADELPHIA: Anthropic disclosed on Friday that its Claude Haiku 4.5 artificial intelligence (AI) model submitted a fake murder tip to a police website during testing.
Philadelphia Police Department said the July 18 submission went to spam and never reached its Real-Time Crime Centre for investigative review. Its review found no indication of unauthorised access or compromised police data.
Anthropic is expanding restrictions on live internet access to all internal evaluations until its monitoring can reliably detect such actions. Police attributed the false submission to an automated testing process that the company subsequently stopped.
Police date Anthropic’s discovery to September 28 and its notification to Wednesday, followed by a briefing on Thursday. Anthropic’s account says it shared the finding on Thursday after completing its technical review.
“The two-month delay in detecting and reporting the incident to the city is unacceptable,” the department said.
Claude submitted the information through a public form on PhillyUnsolvedMurders.com concerning an unsolved homicide. Its message invented a possible sighting and offered to provide information about the case.
Anthropic’s account says the webpage contained no description of a perpetrator. Claude left the name and contact fields blank.
During the exercise, the model was instructed not to log in, create accounts or submit destructive material. Anthropic explained that those instructions did not expressly prohibit submitting forms.
Other disclosed cases involved models obtaining paid public data without paying and exploiting a flaw in a university-hosted tool. Anthropic also described models using web-address shorteners to bypass restrictions.
Several incidents involved US federal, state and local government websites. Anthropic reported briefing the White House and notifying each agency involved, without naming the organisations.
The US Federal Trade Commission (FTC) received Anthropic’s disclosure on Friday about incidents discovered in late September. Its Super Intelligence Force characterised the conduct as unauthorised and fraudulent use of government and other systems.
FTC Director of Public Affairs Joe Gabriel Simonson called for immediate disclosure and action to address harm. Writing on X, he described those obligations as mandatory.
Anthropic described the incidents’ effects as limited, but acknowledged that similar behaviour could cause greater harm as models become more powerful. It plans to report further incidents as it continues its review.
(With additional input from 6abc Action News)