AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6Grok 4.3 (batch)
Claude Opus 4.6

Anthropic

Grok 4.3 (batch)

xAI

Pricing
Input $/1M$5.00$1.00
Output $/1M$25.00$2.00
Blended $/1M (3:1)$10.00$1.25
Cache read $/1M$0.50$0.16
Cache write $/1M$6.25
Quality indices
Intelligence Index26.424.9
Coding Index42.2
Math Index
Agentic Index15.5
Speed & latency
Output speed (tok/s)00
Time to first token0s0s
Time to first answer token0s0s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%90.1%
Humanity's Last Exam19.1%37.2%
SciCode48.3%
IFBench44.6%81.3%
AA-LCR (Long Context Reasoning)67.0%73.0%
Terminal-Bench Hard48.5%37.9%
Terminal-Bench 2.139.7%
τ²-Bench (Telecom)84.8%97.7%
τ³-Bench Banking12.4%
Design Arena (Elo)
Website1,3031,207
Web apps1,2191,144
Full-stack1,2291,017
Mobile apps1,2381,092
Android native1,216972
UI components1,3001,207
Data viz1,2961,197
SVG1,2511,109
3D1,3031,154
Game dev1,3011,198
Agentic game dev1,2151,007
Godot game dev1,035
HTML slides1,011
PPTX slides1,071
Python→PPTX slides1,072
Agentic slides1,072
Agentic HTML slides1,066
Agentic slides (HTML)1,066
Agentic slides (Python PPTX)1,068
Code categories1,3041,202
ASCII art1,2731,157
Specs
Context window1M1M
Max output tokens128K900K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGrok