Zhipu AI said it has launched and open-sourced GLM-5.3-Flash, the first native multimodal model in its GLM-5 series. According to Odaily, the company said the model outperforms GLM-5.2 overall, comes close to Claude Opus 4.8 on coding and agent benchmarks, and costs one-tenth as much as GLM-5.2.
Zhipu AI said GLM-5.3-Flash uses a new base model, introduces a hybrid sparse-attention and linear-attention architecture in the GLM main series for the first time, and was pretrained on 30T tokens of multimodal data. The company also said it tested the model anonymously as Ox-Alpha on OpenCode and OpenRouter before the official release to gather user feedback.
Zhipu AI said Ox-Alpha became the most popular model that week and set a new high for calls on both platforms, with all request traffic supported by domestic chips. The company said earlier community tests of Ox Alpha on DeepSWE briefly showed an 80% result based on 10 questions, but the score fell to about 63% after the sample size was expanded. It added that developers can now deploy the open-sourced weights directly with frameworks including vLLM, SGLang and KTransformers, without using the anonymous model API.
