AiCostCompare

GLM 5.3 FlashX

by Z.ai · z-ai/glm-5.3-flashx · #177 cheapest of 422 paid models

Prices updated Sep 19, 2026, 5:37 PM UTC · refreshed hourly

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.37

per 1M tokens

Output

$1.25

per 1M tokens

Blended (3:1)

$0.59

3 input : 1 output

Cache read

$0.075

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for GLM 5.3 FlashX with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.59$0.039
50%$0.479$0.0242−38%
90%$0.391$0.0124−68%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00056$0.56
Document summary8,0001,000$0.00421$4.21
Codebase question (RAG)30,0002,000$0.0136$13.60
Long-context analysis150,0005,000$0.0617$61.75

Specs

Context window

1.05M

tokens

Max output

131K

tokens

Input modalities

text, image, video

Tokenizer

Other