GLM-5.3-Flash: Z.ai Launches a Multimodal Open Model with 1M Tokens Under MIT License

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-08-27 | By: Carlos Montiel | Reading time: ~4 minutes

Z.ai launched GLM-5.3-Flash on August 26, 2026: a mixture-of-experts (MoE) model with open weights under MIT license, a million-token context window, and native image and video input — which, per the company's own benchmarks, lands half a point behind Claude Opus 4.8 on its internal code test, at roughly a tenth of the price.

The Technical Specs

GLM-5.3-Flash has 320 billion total parameters with 18 billion active per token, trained on a 30-trillion-token multimodal corpus. It combines linear attention for local dependencies with sparse attention for relevant global context, and the weights are published on Hugging Face. It's the first natively multimodal entry in the GLM-5 line — before its name was revealed, it circulated unbranded on OpenCode and OpenRouter under the alias "Ox Alpha."

How It Stacks Up Against Other Models

According to Artificial Analysis, GLM-5.3-Flash scores 57 on its intelligence index, well above the 27 median for open models of similar size. On vision benchmarks like OfficeQA Pro it reaches 62.4, beating both Opus 4.8 and DeepSeek-V4-Vision-Exp, though it trails Gemini 3.7 Flash on tests like BabyVision and MVBench. Z.ai reports it beats its predecessor GLM-5.2 on benchmarks and real workloads at roughly a tenth of the cost.

GLM-5.3-Flash — spec sheet Total parameters: 320B (18B active per token, MoE architecture) Context: 1,048,576 tokens Input: text, image, video (native multimodal) License: MIT (open weights, Hugging Face) Intelligence index (Artificial Analysis): 57 (category median: 27) Training corpus: ~30 trillion multimodal tokens
What this means for companies evaluating open models: GLM-5.3-Flash adds pressure to the price/performance race between open-weight models and proprietary frontier models. For companies in Latin America running code, RAG, or agent workloads on a tight budget, a model that gets close to Claude Opus 4.8 on code at a fraction of the cost — and that can also be self-hosted under an MIT license — is a concrete option worth evaluating for tasks that don't need a proprietary model's absolute top-tier performance, especially in scenarios where data sovereignty requires running the model on your own infrastructure.
Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contact · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com