Model Name
zai-org/GLM-5.3-FlashGLM-5.3-Flash
- Type: Generation
- Capabilities:
reasoning - Cache read: $0.03 per 1M input tokens (0.2× standard input price). See prompt caching.
Overview
GLM-5.3-Flash is Z.ai’s natively multimodal model for coding, reasoning, and agentic workflows. It has 320 billion total parameters with 18 billion active per token, combining sparse and linear attention for efficient long-context processing. Trained on a 30-trillion-token multimodal corpus, it improves on GLM-5.2 across coding, tool-use, and general capability benchmarks while requiring substantially less serving compute.
Reasoning efforts
- Supported:
minimal,low,medium,high,xhigh,max
See the reasoning effort guide for request examples.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.15 | $0.03 | $0.50 |
| Async | $0.11 | $0.02 | $0.38 |
| Batch (24h) | $0.08 | $0.02 | $0.25 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩