AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6GPT-5.4
Claude Opus 4.6

Anthropic

GPT-5.4

OpenAI

Pricing
Input $/1M$5.00$2.50
Output $/1M$25.00$15.00
Blended $/1M (3:1)$10.00$5.63
Cache read $/1M$0.50$0.25
Cache write $/1M$6.25
Quality indices
Intelligence Index37.851.4
Coding Index71.1
Math Index
Agentic Index41.1
Speed & latency
Output speed (tok/s)54149
Time to first token1.89s125.86s
Time to first answer token1.89s125.86s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%92.0%
Humanity's Last Exam18.6%41.6%
SciCode45.7%56.6%
IFBench44.6%73.9%
AA-LCR (Long Context Reasoning)58.3%74.0%
Terminal-Bench Hard48.5%57.6%
Terminal-Bench 2.178.3%
τ²-Bench (Telecom)84.8%87.1%
τ³-Bench Banking30.3%
Design Arena (Elo)
Website1,3231,247
Web apps1,2611,122
Full-stack1,2791,076
Mobile apps1,2801,147
Android native1,1951,020
UI components1,3361,281
Data viz1,3221,268
SVG1,2741,241
3D1,3351,160
Game dev1,3371,298
Godot game dev1,135
Code categories1,3271,245
ASCII art1,2981,241
Specs
Context window1M1.05M
Max output tokens128K128K
Input modalitiestext, image, filetext, image, file
TokenizerClaudeGPT