AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Opus 4.6Mercury 2
Claude Opus 4.6

Anthropic

Mercury 2

Inception

Pricing
Input $/1M$5.00$0.25
Output $/1M$25.00$0.75
Blended $/1M (3:1)$10.00$0.375
Cache read $/1M$0.50$0.025
Cache write $/1M$6.25
Quality indices
Intelligence Index26.413.8
Coding Index31.1
Math Index
Agentic Index1.6
Speed & latency
Output speed (tok/s)0607
Time to first token0s7.16s
Time to first answer token0s7.16s
Benchmarks (Artificial Analysis)
GPQA Diamond84.0%77.0%
Humanity's Last Exam19.1%17.1%
SciCode37.7%
IFBench44.6%69.8%
AA-LCR (Long Context Reasoning)67.0%43.7%
Terminal-Bench Hard48.5%26.5%
Terminal-Bench 2.127.3%
τ²-Bench (Telecom)84.8%70.8%
τ³-Bench Banking9.5%
Design Arena (Elo)
Website1,3031,038
Web apps1,219
Full-stack1,229
Mobile apps1,238
Android native1,216
UI components1,300998
Data viz1,2961,015
SVG1,2511,016
3D1,3031,001
Game dev1,3011,008
Agentic game dev1,215
Code categories1,3041,028
ASCII art1,2731,021
Specs
Context window1M128K
Max output tokens128K50K
Input modalitiestext, image, filetext
TokenizerClaudeOther