AiCostCompare

Mercury 2

by Inception · inception/mercury-2 · #97 cheapest of 325 paid models

Prices updated Jul 22, 2026, 1:40 AM UTC · refreshed hourly

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.25

per 1M tokens

Output

$0.75

per 1M tokens

Blended (3:1)

$0.375

3 input : 1 output

Cache read

$0.025

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for Mercury 2 with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.375$0.0263
50%$0.291$0.015−43%
90%$0.223$0.006−77%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00035$0.35
Document summary8,0001,000$0.00275$2.75
Codebase question (RAG)30,0002,000$0.009$9.00
Long-context analysis150,0005,000$0.0412$41.25

Performance

Intelligence Index

21.4

AA composite quality

Coding Index

31.1

Math Index

Agentic Index

9.6

Output speed

1045 tok/s

median

Time to first token

3.40s

median

Time to first answer

3.40s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
GPQA Diamond77.0%
Humanity's Last Exam15.5%
SciCode38.7%
IFBench69.8%
AA-LCR (Long Context Reasoning)36.3%
Terminal-Bench Hard26.5%
Terminal-Bench 2.127.3%
τ²-Bench (Telecom)70.8%
τ³-Bench Banking9.9%

Design Arena Elo

CategoryElo
Website1,014
UI components1,013
Data viz1,011
SVG1,024
3D1,043
Game dev1,036
Code categories1,025
ASCII art1,046

Specs

Context window

128K

tokens

Max output

50K

tokens

Input modalities

text

Tokenizer

Other