AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6Grok 4.6
Claude Opus 4.6

Anthropic

Grok 4.6

xAI

Pricing
Input $/1M$5.00$2.00
Output $/1M$25.00$6.00
Blended $/1M (3:1)$10.00$3.00
Cache read $/1M$0.50$0.50
Cache write $/1M$6.25
Quality indices
Intelligence Index26.444.3
Coding Index76.8
Math Index
Agentic Index53
Speed & latency
Output speed (tok/s)058
Time to first token0s34.50s
Time to first answer token0s34.50s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%94.9%
Humanity's Last Exam19.1%42.9%
SciCode56.5%
IFBench44.6%
AA-LCR (Long Context Reasoning)67.0%80.3%
Terminal-Bench Hard48.5%
Terminal-Bench 2.188.4%
τ²-Bench (Telecom)84.8%
τ³-Bench Banking50.7%
Design Arena (Elo)
Website1,3031,304
Web apps1,2191,264
Full-stack1,2291,274
Mobile apps1,2381,267
Android native1,2161,301
UI components1,3001,305
Data viz1,2961,303
SVG1,2511,263
3D1,3031,305
Game dev1,3011,321
Agentic game dev1,2151,210
HTML slides1,253
Python→PPTX slides1,242
Code categories1,3041,309
ASCII art1,2731,297
Specs
Context window1M500K
Max output tokens128K450K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGrok