AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Sonnet 5Grok 4.3
Claude Sonnet 5

Anthropic

Grok 4.3

xAI

Pricing
Input $/1M$2.00$1.25
Output $/1M$10.00$2.50
Blended $/1M (3:1)$4.00$1.56
Cache read $/1M$0.20$0.20
Cache write $/1M$2.50
Quality indices
Intelligence Index53.437.6
Coding Index71.542.2
Math Index
Agentic Index46.724.1
Speed & latency
Output speed (tok/s)86122
Time to first token93.73s32.04s
Time to first answer token93.73s32.04s
Benchmarks (Artificial Analysis)
GPQA Diamond91.1%90.1%
Humanity's Last Exam39.6%35.0%
SciCode53.6%47.3%
IFBench81.3%
AA-LCR (Long Context Reasoning)70.7%64.3%
Terminal-Bench Hard37.9%
Terminal-Bench 2.180.5%39.7%
τ²-Bench (Telecom)97.7%
τ³-Bench Banking28.2%12.2%
Design Arena (Elo)
Website1,3141,221
Web apps1,3051,179
Full-stack1,2921,070
Mobile apps1,130
Android native1,2551,056
UI components1,3171,236
Data viz1,2811,222
SVG1,2391,128
3D1,3211,182
Game dev1,3451,235
Agentic game dev1,2411,017
Godot game dev1,2571,061
HTML slides1,2331,045
PPTX slides1,071
Python→PPTX slides1,2521,072
Agentic slides1,072
Agentic HTML slides1,066
Agentic slides (HTML)1,066
Agentic slides (Python PPTX)1,068
Code categories1,3121,221
ASCII art1,2301,184
Specs
Context window1M1M
Max output tokens128K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGrok