Z.ai Launches GLM-5.3-Flash With 1M Context and Lower AI Inference Costs

Z.ai Launches GLM-5.3-Flash With 1M Context and Lower AI Inference Costs

Z.ai has released GLM-5.3-Flash, a 320-billion-parameter multimodal model designed to deliver lower-cost inference while retaining performance close to larger frontier systems. The model activates 18 billion parameters at a time, supports a context window of up to one million tokens and was served during pre-release testing entirely on Chinese AI chips.

GLM-5.3-Flash is the first natively multimodal model in Z.ai’s GLM-5 family. The company has released its weights under an MIT license on Hugging Face and made the model available to GLM Coding Plan users.

Artificial Analysis scored GLM-5.3-Flash at 57 on its Intelligence Index at maximum reasoning effort. That puts it three points behind GLM-5.3, which scored 60, while matching GPT-5.6 Terra and Muse Spark 1.2 in the cited evaluation. The more notable difference is cost. Artificial Analysis measured GLM-5.3-Flash at $0.09 per task, compared with $0.68 for GLM-5.3. Z.ai separately cited a discounted figure of $0.045 per task in its own announcement.

API pricing is set at $0.15 per million input tokens and $0.50 per million output tokens. Z.ai said the model costs roughly one-tenth as much as GLM-5.3 while exceeding GLM-5.2 across its cited benchmarks and real-world workloads.

The company also focused heavily on coding and agentic performance. GLM-5.3-Flash scored 63.4 on DeepSWE v1.1 versus 46.2 for GLM-5.2, while its AutomationBench score rose to 48.8 from 26.2. On Z.ai Code Bench v1.0 at maximum effort, it scored 29.0 compared with 29.5 for Claude Opus 4.8. Artificial Analysis found similar strength on agentic workloads. GLM-5.3-Flash posted an Elo score of roughly 1770 on GDPval-AA v2, putting it alongside GLM-5.3 and Grok 4.6 and behind Claude Opus 5. The model was less efficient in how it used tokens, however, with about 90% of its output tokens going toward reasoning in that evaluation.

Z.ai attributed part of the cost reduction to changes in the model’s architecture. GLM-5.3-Flash combines sparse and linear attention and uses Manifold-Constrained Hyper-Connections. Compared with the GLM-4.5 series, the new model reduces the number of active parameters from 32 billion to 18 billion and cuts the layer count from 92 to 45.

The company also introduced a mechanism called IndexPool to reduce the memory and latency demands of its indexing system at long context lengths. Z.ai said GLM-5.3-Flash uses three times less attention compute and 4.4 times less KV cache than GLM-5.3.

Before the formal release, Z.ai tested the model anonymously under the name ox-alpha on OpenCode and OpenRouter. The company said it became the most popular model of the week during that period. That test also served as a large-scale demonstration of Z.ai’s infrastructure. The company said all ox-alpha traffic ran on Chinese AI chips rather than Nvidia hardware. SemiAnalysis reported that the system processed about 100 trillion tokens per day.

Z.ai built a dedicated inference engine on top of SGLang to improve performance on the underlying hardware. Its serving architecture separates multimodal encoding, prompt prefill and token decoding into independently managed stages, allowing each part of the workload to scale separately.

The company said the optimized stack improved end-to-end serving performance by three times compared with its initial implementation on the same chips. Z.ai also said the resulting hardware efficiency and per-token cost were comparable with mainstream Nvidia GPUs.

GLM-5.3 itself played a role in that optimization work. According to Z.ai, an infrastructure agent powered by the larger model assisted engineers with kernel development, performance bottleneck diagnosis and improvements to the serving stack.

The model’s multimodal capabilities are also intended to extend beyond image understanding. Z.ai said GLM-5.3-Flash can use visual feedback while coding, allowing it to inspect rendered interfaces, interact with environments and revise its work based on what it sees. The company is applying the same approach to documents, spreadsheets, presentations, dashboards and other visually structured work.

Users can also access those capabilities through ZCode with Browser Use and Computer Use. For local deployment, Z.ai currently supports SGLang, vLLM and TokenSpeed.

Z.ai is positioning GLM-5.3-Flash around the combination of lower inference costs, long context, multimodal capabilities and deployment on domestic hardware. The release also gives the company a smaller alternative to GLM-5.3 that approaches its larger model on several cited evaluations while requiring substantially less compute.

This analysis is based on reporting from z.ai & the decoder.

Image courtesy of ArkeonTech.

This article was generated with AI assistance and reviewed for accuracy and quality.

Updated Aug 27, 2026

About this article: This article was generated with AI assistance and reviewed by our editorial team to ensure it follows our editorial standards for accuracy and independence. We maintain strict fact-checking protocols and cite all sources.

Word count: 721Reading time: 0 minutes

📧 Stay Updated

Get the latest AI news delivered to your inbox every morning.

AI News Daily

Breaking Intelligence • Since 2023

Join hundreds of thousands of AI professionals who start their day with our curated newsletter. Get breaking news, expert analysis, and exclusive insights.

Stay Ahead of AI

Get the latest AI breakthroughs, tools, and insights delivered to your inbox every week.

Free forever Unsubscribe anytime No spam guarantee

Go Premium

Unlock unlimited AI tools and an ad-free reading experience designed for AI professionals.

• Ad-free experience• Premium AI tools
Start Free Trial

14-day free trial • Cancel anytime
Plus $9/mo • Pro $90/yr (2 months free)

Follow Our Community

ChatAI

Breaking Intelligence

Your daily briefing on what matters in AI. Trusted by developers, researchers, executives, and AI enthusiasts worldwide.

© 2026 ChatAI. All rights reserved.