AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6DeepSeek V4 Pro
Claude Opus 4.6

Anthropic

DeepSeek V4 Pro

DeepSeek

Pricing
Input $/1M$5.00$0.435
Output $/1M$25.00$0.87
Blended $/1M (3:1)$10.00$0.544
Cache read $/1M$0.50$0.0036
Cache write $/1M$6.25
Quality indices
Intelligence Index37.844.3
Coding Index59.4
Math Index
Agentic Index36.4
Speed & latency
Output speed (tok/s)5474
Time to first token1.89s0.99s
Time to first answer token1.89s60.41s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%88.8%
Humanity's Last Exam18.6%35.9%
SciCode45.7%50.0%
IFBench44.6%76.5%
AA-LCR (Long Context Reasoning)58.3%66.3%
Terminal-Bench Hard48.5%46.2%
Terminal-Bench 2.164.0%
τ²-Bench (Telecom)84.8%96.2%
τ³-Bench Banking25.8%
Design Arena (Elo)
Website1,3231,261
Web apps1,2611,005
Full-stack1,279948
Mobile apps1,280
Android native1,195
UI components1,3361,263
Data viz1,3221,212
SVG1,2741,183
3D1,3351,310
Game dev1,3371,288
Godot game dev1,059
Code categories1,3271,272
ASCII art1,2981,194
Specs
Context window1M1.05M
Max output tokens128K384K
Input modalitiestext, image, filetext
TokenizerClaudeDeepSeek