Z.ai launched GLM-5.3-Flash on August 26, 2026: a mixture-of-experts (MoE) model with open weights under MIT license, a million-token context window, and native image and video input — which, per the company's own benchmarks, lands half a point behind Claude Opus 4.8 on its internal code test, at roughly a tenth of the price.
GLM-5.3-Flash has 320 billion total parameters with 18 billion active per token, trained on a 30-trillion-token multimodal corpus. It combines linear attention for local dependencies with sparse attention for relevant global context, and the weights are published on Hugging Face. It's the first natively multimodal entry in the GLM-5 line — before its name was revealed, it circulated unbranded on OpenCode and OpenRouter under the alias "Ox Alpha."
According to Artificial Analysis, GLM-5.3-Flash scores 57 on its intelligence index, well above the 27 median for open models of similar size. On vision benchmarks like OfficeQA Pro it reaches 62.4, beating both Opus 4.8 and DeepSeek-V4-Vision-Exp, though it trails Gemini 3.7 Flash on tests like BabyVision and MVBench. Z.ai reports it beats its predecessor GLM-5.2 on benchmarks and real workloads at roughly a tenth of the cost.
Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.
Contact Carlos Montiel