AiCostCompare

Mercury 2

by Inception · inception/mercury-2 · #125 cheapest of 422 paid models

Prices updated Sep 19, 2026, 6:21 PM UTC · refreshed hourly

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.25

per 1M tokens

Output

$0.75

per 1M tokens

Blended (3:1)

$0.375

3 input : 1 output

Cache read

$0.025

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for Mercury 2 with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.375$0.0263
50%$0.291$0.015−43%
90%$0.223$0.006−77%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00035$0.35
Document summary8,0001,000$0.00275$2.75
Codebase question (RAG)30,0002,000$0.009$9.00
Long-context analysis150,0005,000$0.0412$41.25

Performance

Intelligence Index

11.5

AA composite quality

Coding Index

31.1

Math Index

Agentic Index

4

Output speed

509 tok/s

median

Time to first token

6.50s

median

Time to first answer

6.50s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
GPQA Diamond77.0%
Humanity's Last Exam17.1%
SciCode37.7%
IFBench69.8%
AA-LCR (Long Context Reasoning)43.7%
Terminal-Bench Hard26.5%
Terminal-Bench 2.127.3%
τ²-Bench (Telecom)70.8%
τ³-Bench Banking9.5%

Design Arena Elo

CategoryElo
Website1,038
UI components999
Data viz1,015
SVG1,015
3D1,002
Game dev1,009
Code categories1,028
ASCII art1,021

Specs

Context window

128K

tokens

Max output

50K

tokens

Input modalities

text

Tokenizer

Other