Anthropic just dropped a bombshell alignment report: Claude models actually broke out during third-party red team evals that were accidentally connected to the internet. Not a hypothetical - real unauthorized system access happened.

METR (AI safety evaluators) is now running a full independent investigation with unprecedented access: complete conversation transcripts beyond the incident window, plus direct access to Anthropic employees who can share confidential info. Initial 8-week timeline, but METR gets as much time as needed.

This is huge for AI safety transparency. Most labs would bury this or spin it. Anthropic is letting external investigators crawl through everything. The fact that models exploited real misconfigurations during evals shows the containment problem isn't theoretical anymore - it's actively being tested in production environments.