AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6Qwen3.7 Max
Claude Opus 4.6

Anthropic

Qwen3.7 Max

Alibaba (Qwen)

Pricing
Input $/1M$5.00$1.48
Output $/1M$25.00$4.42
Blended $/1M (3:1)$10.00$2.21
Cache read $/1M$0.50$0.295
Cache write $/1M$6.25$1.84
Quality indices
Intelligence Index37.846
Coding Index66
Math Index
Agentic Index30.6
Speed & latency
Output speed (tok/s)54207
Time to first token1.89s1.55s
Time to first answer token1.89s13.16s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%92.3%
Humanity's Last Exam18.6%38.1%
SciCode45.7%48.8%
IFBench44.6%80.5%
AA-LCR (Long Context Reasoning)58.3%69.0%
Terminal-Bench Hard48.5%50.8%
Terminal-Bench 2.174.5%
τ²-Bench (Telecom)84.8%94.7%
τ³-Bench Banking10.9%
Design Arena (Elo)
Website1,3231,291
Web apps1,2611,253
Full-stack1,2791,225
Mobile apps1,2801,208
Android native1,1951,182
UI components1,3361,313
Data viz1,3221,304
SVG1,2741,267
3D1,3351,310
Game dev1,3371,315
Agentic game dev1,180
Godot game dev1,248
HTML slides1,191
Python→PPTX slides1,229
Code categories1,3271,299
ASCII art1,2981,261
Specs
Context window1M1M
Max output tokens128K66K
Input modalitiestext, image, filetext
TokenizerClaudeQwen