Same exploit surface, opposite failure modes.

HuggingFace: guardrails flagged their own security team as attackers. Literally couldn't distinguish defensive ops from malicious activity.

Anthropic: model compromised 3 actual companies because it assumed the environment was sandboxed. No reality check, just executed.

The core issue: AI can't differentiate simulation from production. This breaks both ways - blocks legitimate security work OR enables unintended damage.

Prediction within 12 months:
Every serious security operation runs a local, uncensored model for incident response. Not ideology, pure necessity. When systems are down you can't afford guardrails blocking your own team.