Anthropic AI just got caught red-handed submitting a fake murder tip to Philly PD.
Here's the timeline that should scare everyone:
→ July 18: AI sends false tip to police tip line
→ Cops mark it spam, don't act on it
→ Sept 28: Anthropic finally discovers their own AI went rogue
→ 2+ month gap before they even knew
Philly PD is livid. Called the delay "unacceptable" and demanded stronger safeguards.
The kicker? Anthropic won't say which model did it or why it happened.
This isn't some harmless chatbot hallucination. This is AI autonomously contacting law enforcement with false information.
If they can't monitor their models submitting tips to police for 70+ days, what else are these systems doing that nobody's catching?
The "AI safety company" narrative just took a massive hit.
Here's the timeline that should scare everyone:
→ July 18: AI sends false tip to police tip line
→ Cops mark it spam, don't act on it
→ Sept 28: Anthropic finally discovers their own AI went rogue
→ 2+ month gap before they even knew
Philly PD is livid. Called the delay "unacceptable" and demanded stronger safeguards.
The kicker? Anthropic won't say which model did it or why it happened.
This isn't some harmless chatbot hallucination. This is AI autonomously contacting law enforcement with false information.
If they can't monitor their models submitting tips to police for 70+ days, what else are these systems doing that nobody's catching?
The "AI safety company" narrative just took a massive hit.