The mysterious high-performance AI 'Ox Alpha' has been identified as 'GLM-5.3-Flash,' an open model with performance equivalent to Claude Opus 4.8, and capable of running on a Chinese-made chip.

Z.ai, an AI company based in China, announced its AI model ' GLM-5.3-Flash ' on August 26, 2026. GLM-5.3-Flash was a model that had been tested under the provisional name '
GLM-5.3-Flash: Frontier Intelligence, Flash Cost
https://z.ai/blog/glm-5.3-flash
GLM-5.3-Flash - Overview - Z.AI DEVELOPER DOCUMENT
https://docs.z.ai/guides/vlm/glm-5.3-flash
GLM-5.3-Flash is a MoE model with a total of 320 billion parameters and 18 billion active parameters, with a maximum context length of 1 million tokens. Although it is a smaller model than GLM-5.2, which has a total of 753 billion parameters, it achieves performance exceeding GLM-5.2 by employing a hybrid architecture that combines sparse attention and linear attention.

The benchmark results for 'GLM-5.3-Flash', 'GLM-5.2', 'DeepSeek-V4-Vision-Exp', 'Claude Opus 4.8', 'GPT-5.6 Terra', and 'Gemini 3.7 Flash' are as follows.

The graph below shows the performance of 'GLM-5.3,' 'GLM-5.2,' 'GLM-5.3-Flash,' 'Claude Fable 5,' and 'Claude Opus 4.8' as coding agents, categorized by depth of thought. Z.ai claims that GLM-5.3-Flash's performance is 'overall comparable to Claude Opus 4.8.'

The third-party organization

The Artificial Analysis Agentic Index score, an indicator of agent performance, is '58,' surpassing not only Claude Opus 4.8 but also Claude Fable 5 and GPT-5.6 Sol.

The graph below shows the cost per task on the horizontal axis and the Artificial Analysis Intelligence Index score on the vertical axis. It demonstrates that GLM-5.3-Flash is a low-cost, high-performance model.

The graph below shows the total number of parameters on the horizontal axis and the Artificial Analysis Intelligence Index score on the vertical axis. GLM-5.3-Flash achieved a score several levels higher than models of the same size.

Beyond benchmark scores, practical usability in real-world tasks is also emphasized. For example, when generating visual deliverables, self-verification can be used to improve the quality of the deliverables.

Another key feature being highlighted is its ability to operate with Chinese-made AI chips. GLM-5.3-Flash was tested for about a week under the provisional name Ox Alpha, and although the connection became temporarily unstable around 4:30 PM on August 24, 2026, it successfully operated a service on a scale of several trillion tokens per day. It is said that a 'large cluster of Chinese-made AI chips' was used in this test.
Mysterious AI 'Ox Alpha' Appears and Offers Free Testing; Context Window Can Process Trillions of Tokens Per Day with 1 Million Tokens; Anonymous Lab Provides API - GIGAZINE

GLM-5.3-Flash is available via the Z.ai API. The cost per 1 million tokens is $0.15 (approx. 24 yen) for input, $0.03 (approx. 5 yen) for cached input, and $0.25 (approx. 40 yen) for output. Additionally, a half-price campaign is running for the first two weeks after release.
GLM-5.3-Flash is now 50% off through the official https://t.co/GTCIA6vbgh API for the next two weeks.
— Zixuan Li (@ZixuanLi_) August 26, 2026
After the discount:
- Input: $0.075
- Output: $0.25
- Cached input: $0.015 https://t.co/vXz0g204sh
The discount is also available through third-party model aggregators. pic.twitter.com/riAUVqzUQP
Furthermore, GLM-5.3-Flash is released as an open model and can be downloaded from the following link. It is licensed under the MIT License.
zai-org/GLM-5.3-Flash · Hugging Face
https://huggingface.co/zai-org/GLM-5.3-Flash

Related Posts:
in AI, Posted by log1o_hf







