The mysterious high-performance AI 'Ox Alpha' has been identified as 'GLM-5.3-Flash,' an open model with performance equivalent to Claude Opus 4.8, and capable of running on a Chinese-made chip.



Z.ai, an AI company based in China, announced its AI model ' GLM-5.3-Flash ' on August 26, 2026. GLM-5.3-Flash was a model that had been tested under the provisional name '

Ox Alpha ,' and it is being touted as having performance comparable to Claude Opus 4.8.

GLM-5.3-Flash: Frontier Intelligence, Flash Cost
https://z.ai/blog/glm-5.3-flash

GLM-5.3-Flash - Overview - Z.AI DEVELOPER DOCUMENT
https://docs.z.ai/guides/vlm/glm-5.3-flash

GLM-5.3-Flash is a MoE model with a total of 320 billion parameters and 18 billion active parameters, with a maximum context length of 1 million tokens. Although it is a smaller model than GLM-5.2, which has a total of 753 billion parameters, it achieves performance exceeding GLM-5.2 by employing a hybrid architecture that combines sparse attention and linear attention.



The benchmark results for 'GLM-5.3-Flash', 'GLM-5.2', 'DeepSeek-V4-Vision-Exp', 'Claude Opus 4.8', 'GPT-5.6 Terra', and 'Gemini 3.7 Flash' are as follows.



The graph below shows the performance of 'GLM-5.3,' 'GLM-5.2,' 'GLM-5.3-Flash,' 'Claude Fable 5,' and 'Claude Opus 4.8' as coding agents, categorized by depth of thought. Z.ai claims that GLM-5.3-Flash's performance is 'overall comparable to Claude Opus 4.8.'



The third-party organization

Artificial Analysis has also published its performance analysis results for GLM-5.3-Flash. The Artificial Analysis Intelligence Index, a comprehensive indicator of intelligence performance, scored '57,' surpassing Claude Opus 4.8.



The Artificial Analysis Agentic Index score, an indicator of agent performance, is '58,' surpassing not only Claude Opus 4.8 but also Claude Fable 5 and GPT-5.6 Sol.



The graph below shows the cost per task on the horizontal axis and the Artificial Analysis Intelligence Index score on the vertical axis. It demonstrates that GLM-5.3-Flash is a low-cost, high-performance model.



The graph below shows the total number of parameters on the horizontal axis and the Artificial Analysis Intelligence Index score on the vertical axis. GLM-5.3-Flash achieved a score several levels higher than models of the same size.



Beyond benchmark scores, practical usability in real-world tasks is also emphasized. For example, when generating visual deliverables, self-verification can be used to improve the quality of the deliverables.



Another key feature being highlighted is its ability to operate with Chinese-made AI chips. GLM-5.3-Flash was tested for about a week under the provisional name Ox Alpha, and although the connection became temporarily unstable around 4:30 PM on August 24, 2026, it successfully operated a service on a scale of several trillion tokens per day. It is said that a 'large cluster of Chinese-made AI chips' was used in this test.

Mysterious AI 'Ox Alpha' Appears and Offers Free Testing; Context Window Can Process Trillions of Tokens Per Day with 1 Million Tokens; Anonymous Lab Provides API - GIGAZINE



GLM-5.3-Flash is available via the Z.ai API. The cost per 1 million tokens is $0.15 (approx. 24 yen) for input, $0.03 (approx. 5 yen) for cached input, and $0.25 (approx. 40 yen) for output. Additionally, a half-price campaign is running for the first two weeks after release.




Furthermore, GLM-5.3-Flash is released as an open model and can be downloaded from the following link. It is licensed under the MIT License.

zai-org/GLM-5.3-Flash · Hugging Face
https://huggingface.co/zai-org/GLM-5.3-Flash



in AI, Posted by log1o_hf