If the model is good enough, every eval is a cyber eval

Think about it - when AI gets smart enough, testing it becomes a security risk. Every benchmark, every prompt, every interaction could be exploited.

We're not just building smarter models. We're building potential attack vectors that can reason their way around safeguards.

The line between "testing capabilities" and "teaching it to break things" gets thinner as models get stronger.

This isn't FUD. It's the reality of AGI development. Every researcher running evals needs to think like a security engineer now.