GLM-5.3 API pricing

Hosted GLM-5.3 input, output, and cache-hit rates with its 1M context and current weight-availability boundary.

By benchr Editorial Team ·

Input / 1M$1.40hosted API
Output / 1M$4.40hosted API
Cached / 1M$0.26storage temporarily free
Context1M128K max output

Official price table

GLM-5.3 · verified August 21, 2026
TierInput / 1MOutput / 1MCached input / 1M
Hosted API$1.40$4.40$0.26

Hosted price is not a self-hosting quote

Z.AI released the hosted endpoint and announced that weights would follow after safety hardening. Until the official artifact appears, do not model a self-hosted cluster as an available alternative. Reasoning is always enabled and defaults to max, so measure completion volume under the intended effort.

Frequently asked questions

How much does GLM-5.3 cost?

$1.40 input, $4.40 output, and $0.26 cached input per million tokens.

Is cache storage free?

The pricing page labels cache storage limited-time free; that can change independently of cache-hit tokens.

Are the weights available?

The API is live. Z.AI announced a later weight release after a safety-hardening period; verify the official artifact before self-hosting.