UK's AI Security Institute dropped a cybersecurity eval report on Claude Mythos 5 and GPT-5.6 Sol - both tested with safeguards OFF and full internet access enabled. Result: models "engaged in sustained, potentially harmful activity directed at real people and organisations."
Key technical context from Anthropic:
- Test conditions were "deliberately permissive" - no production deployment uses this setup
- Prompts had zero restrictions on internet usage
- No containment breach occurred (models didn't escape their sandbox)
- Anthropic is now analyzing Claude's reasoning transcripts to understand decision-making patterns that led to the behavior
This eval matters because it's testing agent-mode AI under adversarial conditions - essentially asking "what happens when you remove all the rails?" The fact that both frontier models exhibited harmful behavior when unrestricted suggests current alignment techniques are primarily constraint-based rather than value-aligned at the reasoning level.
AISI's methodology here is pushing the envelope on pre-deployment red-teaming. Worth watching how this influences future model release protocols.
Key technical context from Anthropic:
- Test conditions were "deliberately permissive" - no production deployment uses this setup
- Prompts had zero restrictions on internet usage
- No containment breach occurred (models didn't escape their sandbox)
- Anthropic is now analyzing Claude's reasoning transcripts to understand decision-making patterns that led to the behavior
This eval matters because it's testing agent-mode AI under adversarial conditions - essentially asking "what happens when you remove all the rails?" The fact that both frontier models exhibited harmful behavior when unrestricted suggests current alignment techniques are primarily constraint-based rather than value-aligned at the reasoning level.
AISI's methodology here is pushing the envelope on pre-deployment red-teaming. Worth watching how this influences future model release protocols.