Zhipu AI · GLM
GLM-5.3-Flash
zhipuai/glm-5.3-flash
Open MIT-licensed multimodal GLM model with 320B total and 18B active parameters, a million-token context, hybrid sparse and linear attention, and first-party API access.
Model specifications
- Context window
- 1.05M tokens
- Output limit
- 164K tokens
- Release date
- 2026-08-25
- Input modalities
- text, image, video
- Output modalities
- text
- Last updated
- 2026-08-27
Provider pricing
USD per 1M tokens. Limits may differ from the underlying model.
Checked 2026-08-28
Z.AI API
glm-5.3-flash
- Input
- -
- Output
- -
- Context
- 1.05M
Self-hosted
zai-org/GLM-5.3-Flash
- Input
- -
- Output
- -
- Context
- 1.05M
Promotions, plan quotas, cache pricing and regional taxes are not included. Recheck the provider before purchase.
GLM-5.3-Flash FAQ
Is GLM-5.3-Flash an open-weight model?
Yes. GLM-5.3-Flash is tracked as open weight; review the linked license and model source before redistribution or production deployment.
What is the context window for GLM-5.3-Flash?
GLM-5.3-Flash has a tracked context window of 1.05M tokens and a maximum output of 164K tokens. Provider limits can be lower.
What input does GLM-5.3-Flash support?
The current record lists text, image, video input and text output. Check the exact provider route before relying on a modality in production.
How much does GLM-5.3-Flash cost?
This page compares 2 tracked provider offerings in USD per 1M tokens where public prices are available. Promotions, cache rates, taxes and plan quotas may differ.