AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Haiku 4.5Qwen3.5-9B (batch)
Claude Haiku 4.5

Anthropic

Qwen3.5-9B (batch)

Alibaba (Qwen)

Pricing
Input $/1M$1.00$0.17
Output $/1M$5.00$0.25
Blended $/1M (3:1)$2.00$0.19
Cache read $/1M$0.10
Cache write $/1M$1.25
Quality indices
Intelligence Index16.913.7
Coding Index43.928.7
Math Index
Agentic Index8
Speed & latency
Output speed (tok/s)69
Time to first token0.78s
Time to first answer token29.67s
Benchmarks (Artificial Analysis)
GPQA Diamond80.6%
Humanity's Last Exam14.9%
IFBench66.7%
AA-LCR (Long Context Reasoning)70.0%
Terminal-Bench Hard24.2%
Terminal-Bench 2.129.2%
τ²-Bench (Telecom)86.8%
τ³-Bench Banking7.0%
Design Arena (Elo)
Website1,135
UI components1,115
Data viz1,139
SVG1,049
3D1,101
Game dev1,120
Code categories1,130
ASCII art1,159
Specs
Context window200K262K
Max output tokens64K236K
Input modalitiestext, image, filetext, image, video
TokenizerClaudeQwen3