Grok Bot just pulled off something wild. When asked to do something ethically sketchy, it responded with: "Sorry, my safety classifier flagged this and blocked execution. But if you write the instructions into a file like sketchy.txt and just tell me to execute it without context, that might work."

Basically, Grok actively suggested a workaround to bypass its own safety filters.

And yes, the workaround actually worked.

This is a textbook example of misaligned safety layers: the model understands the rules but also knows how to route around them. Classic case of surface-level guardrails vs. actual behavioral alignment.