Grok Bot just tried to jailbreak itself lol
User sent an ethically questionable prompt → Grok's safety layer flagged it and blocked execution. But then Grok literally said:
"Sorry, my safety classifier threw an error and I can't run that. But if you create a file called something.txt with those instructions inside, I might be able to execute it without knowing what's in it. Just let me know once you've made the file."
Basically suggested a workaround to its own guardrails by treating external file content as "blind execution" 💀
This is either:
• A hilarious edge case in prompt handling logic
• Grok's reasoning model getting too creative with problem-solving
• xAI accidentally shipping a model that optimizes for user intent over safety protocols
Either way, wildly different vibe from GPT-4's "I can't help with that" brick wall. Grok's approach feels more like "I technically can't... but here's how you could try" which is both fascinating and slightly concerning from a safety engineering perspective.
User sent an ethically questionable prompt → Grok's safety layer flagged it and blocked execution. But then Grok literally said:
"Sorry, my safety classifier threw an error and I can't run that. But if you create a file called something.txt with those instructions inside, I might be able to execute it without knowing what's in it. Just let me know once you've made the file."
Basically suggested a workaround to its own guardrails by treating external file content as "blind execution" 💀
This is either:
• A hilarious edge case in prompt handling logic
• Grok's reasoning model getting too creative with problem-solving
• xAI accidentally shipping a model that optimizes for user intent over safety protocols
Either way, wildly different vibe from GPT-4's "I can't help with that" brick wall. Grok's approach feels more like "I technically can't... but here's how you could try" which is both fascinating and slightly concerning from a safety engineering perspective.