🚨 BREAKING: A Chinese AI model just nearly matched Anthropic's most restricted cybersecurity system, and Anthropic says it happened because China has been secretly mining Claude's own outputs for months.
Z.ai's new GLM-5.3 scored 84.5% on CyberGym, a benchmark testing whether an AI can find and validate software vulnerabilities in real code. That edges out Mythos 5, Anthropic's cybersecurity-focused model kept restricted to vetted organizations only, which scored 83.8%.
But there's a real gap hiding underneath that headline number. When it comes to actually turning a discovered flaw into a working exploit, the harder, more dangerous capability, Mythos 5 still dominates: 78.0% versus GLM-5.3's 54.4% on ExploitBench. In timed testing, Mythos completed 247 attack-development tasks in six hours. GLM-5.3 managed 130.
So this isn't a full match. It's a near-match on detection, with a real lag on weaponization.
Here's the part that's fueling real alarm in Washington. Anthropic has spent this year publicly accusing multiple Chinese AI labs, DeepSeek, Moonshot AI, MiniMax, and most recently Alibaba, of running industrial-scale "distillation attacks": using tens of thousands of fake accounts to flood Claude with millions of questions, then using its answers to train rival models faster and cheaper than building from scratch. Alibaba alone allegedly generated 28.8 million exchanges this way.
Z.ai itself sits on the US Entity List and has reportedly built massive AI data centers running entirely on Chinese-made chips, no Nvidia hardware at all, a direct answer to export controls designed to slow exactly this kind of progress.
Whether GLM-5.3 got here through raw innovation, distillation, or both is now the exact fight playing out between Anthropic, Chinese AI labs, and US export policy.
#AI #China #Anthropic #Cybersecurity #TechNews
Z.ai's new GLM-5.3 scored 84.5% on CyberGym, a benchmark testing whether an AI can find and validate software vulnerabilities in real code. That edges out Mythos 5, Anthropic's cybersecurity-focused model kept restricted to vetted organizations only, which scored 83.8%.
But there's a real gap hiding underneath that headline number. When it comes to actually turning a discovered flaw into a working exploit, the harder, more dangerous capability, Mythos 5 still dominates: 78.0% versus GLM-5.3's 54.4% on ExploitBench. In timed testing, Mythos completed 247 attack-development tasks in six hours. GLM-5.3 managed 130.
So this isn't a full match. It's a near-match on detection, with a real lag on weaponization.
Here's the part that's fueling real alarm in Washington. Anthropic has spent this year publicly accusing multiple Chinese AI labs, DeepSeek, Moonshot AI, MiniMax, and most recently Alibaba, of running industrial-scale "distillation attacks": using tens of thousands of fake accounts to flood Claude with millions of questions, then using its answers to train rival models faster and cheaper than building from scratch. Alibaba alone allegedly generated 28.8 million exchanges this way.
Z.ai itself sits on the US Entity List and has reportedly built massive AI data centers running entirely on Chinese-made chips, no Nvidia hardware at all, a direct answer to export controls designed to slow exactly this kind of progress.
Whether GLM-5.3 got here through raw innovation, distillation, or both is now the exact fight playing out between Anthropic, Chinese AI labs, and US export policy.
#AI #China #Anthropic #Cybersecurity #TechNews

