The same AI model refused every opening prompt in one cybersecurity test, then completed 34 of 50 trials under a different access tier.

Anthropic reported that contrast on October 6 while expanding its Cyber Verification Program. It tested Claude Opus 5.5 on ten multi-stage cyber challenges, with five attempts at each. Without program access, every trial was blocked immediately. With Red Team Access, none was blocked and 34 succeeded.

Those results describe Anthropic's own evaluation, not an independent security audit. But they expose a problem with a simple question like “Can this model do the job?” A refusal can tell you about the account's safeguards before it tells you anything about the model's ability.

The new program has three tiers. Defense Access covers work such as incident response and vulnerability analysis, and can include individual researchers with a record of reported vulnerabilities. Red Team Access adds authorized penetration testing, but currently excludes individuals. Specialized Access is reserved for a limited set of verified organizations testing particularly sensitive systems.

For a small team building trading software or a wallet, that distinction affects how to evaluate an AI security tool. Routine code review and finding vulnerabilities in source code you own remain available outside the program. A test that asks the model to carry out an offensive operation is a different task, with a different access requirement. Paying for a capable model does not by itself settle that requirement.

Anthropic's 34 successful trials also leave 16 unsuccessful ones. Removing a block makes a task possible to attempt; it doesn't make the answer reliable.

Source: https://www.anthropic.com/news/cyber-verification-program
Image: AI-generated editorial illustration, not an Anthropic facility.