Z.ai Unveils GLM-5.3-Flash: A 320B-A18B Multimodal MoE Model with 1M-Token Context

Z.ai has officially released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series and the most cost-effective coding model the lab has shipped to date. Built as a mixture-of-experts (MoE) architecture with 320B total parameters and 18B active parameters per token, it supports a 1,048,576-token context window and processes both image and video inputs. The model is available under an MIT license, with weights hosted on Hugging Face.


According to Z.ai, GLM-5.3-Flash outperforms its predecessor, GLM-5.2, across both benchmark suites and real-world workloads at approximately one-tenth the inference cost. Internally, it lands within half a point of Claude Opus 4.8 on coding evaluations. Interestingly, the model was quietly tested during its first week under the alias "Ox Alpha" on OpenCode and OpenRouter, running entirely on domestically produced Chinese AI chips—underscoring a shift toward hardware independence.


Deployment Options


For those eager to integrate the model, GLM-5.3-Flash is deployable through two primary channels:


  1. Open Weights: The full model weights are available on Hugging Face under the permissive MIT license, allowing for self-hosting and fine-tuning.
  2. Hosted API: Z.ai offers a fully managed API that is already priced and operational, enabling immediate integration without infrastructure setup.

  3. This dual-track approach aligns with the broader 2026 trend of open-weight models bridging the gap between research transparency and commercial viability. Whether you are building agentic AI systems, analyzing video content, or tackling complex coding tasks, GLM-5.3-Flash aims to deliver enterprise-grade performance with unprecedented cost efficiency.

    via MarkTechPost

Related