Artificial Analysis scored GLM-5.3-Flash at 57 on its Intelligence Index at maximum reasoning effort. That puts it three points behind GLM-5.3, which scored 60, while matching GPT-5.6 Terra and Muse Spark 1.2 in the cited evaluation. The more notable difference is cost. Artificial Analysis measured GLM-5.3-Flash at $0.09 per task, compared with $0.68 for GLM-5.3. Z.ai separately cited a discounted figure of $0.045 per task in its own announcement.
API pricing is set at $0.15 per million input tokens and $0.50 per million output tokens. Z.ai said the model costs roughly one-tenth as much as GLM-5.3 while exceeding GLM-5.2 across its cited benchmarks and real-world workloads.
The company also focused heavily on coding and agentic performance. GLM-5.3-Flash scored 63.4 on DeepSWE v1.1 versus 46.2 for GLM-5.2, while its AutomationBench score rose to 48.8 from 26.2. On Z.ai Code Bench v1.0 at maximum effort, it scored 29.0 compared with 29.5 for Claude Opus 4.8. Artificial Analysis found similar strength on agentic workloads. GLM-5.3-Flash posted an Elo score of roughly 1770 on GDPval-AA v2, putting it alongside GLM-5.3 and Grok 4.6 and behind Claude Opus 5. The model was less efficient in how it used tokens, however, with about 90% of its output tokens going toward reasoning in that evaluation.
Z.ai attributed part of the cost reduction to changes in the model’s architecture. GLM-5.3-Flash combines sparse and linear attention and uses Manifold-Constrained Hyper-Connections. Compared with the GLM-4.5 series, the new model reduces the number of active parameters from 32 billion to 18 billion and cuts the layer count from 92 to 45.
The company also introduced a mechanism called IndexPool to reduce the memory and latency demands of its indexing system at long context lengths. Z.ai said GLM-5.3-Flash uses three times less attention compute and 4.4 times less KV cache than GLM-5.3.
Before the formal release, Z.ai tested the model anonymously under the name ox-alpha on OpenCode and OpenRouter. The company said it became the most popular model of the week during that period. That test also served as a large-scale demonstration of Z.ai’s infrastructure. The company said all ox-alpha traffic ran on Chinese AI chips rather than Nvidia hardware. SemiAnalysis reported that the system processed about 100 trillion tokens per day.
Z.ai built a dedicated inference engine on top of SGLang to improve performance on the underlying hardware. Its serving architecture separates multimodal encoding, prompt prefill and token decoding into independently managed stages, allowing each part of the workload to scale separately.
The company said the optimized stack improved end-to-end serving performance by three times compared with its initial implementation on the same chips. Z.ai also said the resulting hardware efficiency and per-token cost were comparable with mainstream Nvidia GPUs.
GLM-5.3 itself played a role in that optimization work. According to Z.ai, an infrastructure agent powered by the larger model assisted engineers with kernel development, performance bottleneck diagnosis and improvements to the serving stack.
The model’s multimodal capabilities are also intended to extend beyond image understanding. Z.ai said GLM-5.3-Flash can use visual feedback while coding, allowing it to inspect rendered interfaces, interact with environments and revise its work based on what it sees. The company is applying the same approach to documents, spreadsheets, presentations, dashboards and other visually structured work.
Users can also access those capabilities through ZCode with Browser Use and Computer Use. For local deployment, Z.ai currently supports SGLang, vLLM and TokenSpeed.
Z.ai is positioning GLM-5.3-Flash around the combination of lower inference costs, long context, multimodal capabilities and deployment on domestic hardware. The release also gives the company a smaller alternative to GLM-5.3 that approaches its larger model on several cited evaluations while requiring substantially less compute.
This analysis is based on reporting from z.ai & the decoder.
Image courtesy of ArkeonTech.
This article was generated with AI assistance and reviewed for accuracy and quality.