AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6GPT-5.3-Codex
Claude Opus 4.6

Anthropic

GPT-5.3-Codex

OpenAI

Pricing
Input $/1M$5.00$1.75
Output $/1M$25.00$14.00
Blended $/1M (3:1)$10.00$4.81
Cache read $/1M$0.50$0.175
Cache write $/1M$6.25
Quality indices
Intelligence Index37.844.3
Coding Index
Math Index
Agentic Index
Speed & latency
Output speed (tok/s)54126
Time to first token1.89s51.38s
Time to first answer token1.89s51.38s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%91.5%
Humanity's Last Exam18.6%39.9%
SciCode45.7%53.2%
IFBench44.6%75.4%
AA-LCR (Long Context Reasoning)58.3%74.0%
Terminal-Bench Hard48.5%53.0%
τ²-Bench (Telecom)84.8%86.0%
Design Arena (Elo)
Website1,3231,190
Web apps1,2611,103
Full-stack1,2791,041
Mobile apps1,2801,127
Android native1,1951,084
UI components1,3361,179
Data viz1,3221,199
SVG1,2741,176
3D1,3351,067
Game dev1,3371,219
Godot game dev1,124
Code categories1,3271,178
ASCII art1,2981,193
Specs
Context window1M400K
Max output tokens128K128K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGPT