GLM 5.2 API pricing

Z.ai's flagship open-weight GLM model — sold at the same $1.40/$4.40 list price by Fireworks, Together, and Mistral. Prices below are Fireworks AI's — open-weight models are priced by each host, not the model's creator.

Tool callingReasoning

Input

$1.40

per 1M tokens

Cached input

$0.14

per 1M tokens

Output

$4.40

per 1M tokens

Context window

1M

max output 128K

Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Hand-checked against the provider's official pricing page on Aug 20, 2026. Spot an error? Tell us.

Where to buy it

Each host prices the model independently — same weights, different bill.

HostInput / 1MOutput / 1MContext
Fireworks AI$1.40$4.401M

What would GLM 5.2 cost you?

Enter your workload — add other models to compare the same traffic across providers.

Your monthly workload

Cache hit rate = share of input tokens served from the provider's prompt cache, billed at the model's cached-input price. Estimates use standard per-token rates — batch discounts, service tiers, and long-context surcharges are not applied.

GLM 5.2 (Fireworks AI)
ModelInputCached inputOutputPer requestMonthly
GLM 5.2 (Fireworks AI)$840$56.00$1,100$0.00399$1,996
Rates used: GLM 5.2 (Fireworks AI) $1.40 in / $4.40 out

Tracking GLM 5.2 in production?

Marginal prices every call your app actually makes — cached tokens, dated snapshots, model switches included — and slices spend by customer or feature. One track() call to integrate.