AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Haiku 4.5Grok 4.20
Claude Haiku 4.5

Anthropic

Grok 4.20

xAI

Pricing
Input $/1M$1.00$1.25
Output $/1M$5.00$2.50
Blended $/1M (3:1)$2.00$1.56
Cache read $/1M$0.10$0.20
Cache write $/1M$1.25
Quality indices
Intelligence Index29.637
Coding Index43.9
Math Index
Agentic Index16.4
Speed & latency
Output speed (tok/s)230
Time to first token16.10s
Time to first answer token16.10s
Benchmarks (Artificial Analysis)
GPQA Diamond91.1%
Humanity's Last Exam32.2%
SciCode45.6%
IFBench81.2%
AA-LCR (Long Context Reasoning)58.0%
Terminal-Bench Hard37.9%
τ²-Bench (Telecom)93.0%
Design Arena (Elo)
Website1,1491,255
Web apps1,200
Full-stack1,111
Mobile apps1,181
Android native1,154
UI components1,1421,245
Data viz1,1601,241
SVG1,0721,211
3D1,1311,249
Game dev1,1551,255
Godot game dev1,089
HTML slides1,180
Code categories1,1481,252
ASCII art1,1821,219
Specs
Context window200K2M
Max output tokens64K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGrok