Mercury 2.5
by Inception · inception/mercury-2.5 · #21 cheapest of 422 paid models
Prices updated Sep 19, 2026, 6:22 PM UTC · refreshed hourly
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol
Pricing
Input
$0.04
per 1M tokens
Output
$0.15
per 1M tokens
Blended (3:1)
$0.068
3 input : 1 output
Cache read
$0.0040
per 1M cached tokens
Cache write
—
per 1M tokens
Cache & batch economics
Effective prices for Mercury 2.5 with prompt caching.
| Cache hit rate | Effective blended $/1M | Example request* | vs no cache |
|---|---|---|---|
| 0% | $0.068 | $0.00423 | — |
| 50% | $0.054 | $0.00243 | −43% |
| 90% | $0.043 | $0.00099 | −77% |
*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.
What a request costs
| Workload | Input tokens | Output tokens | Cost / request | Cost / 1K requests |
|---|---|---|---|---|
| Short chat message | 500 | 300 | $6.50e-5 | $0.065 |
| Document summary | 8,000 | 1,000 | $0.00047 | $0.47 |
| Codebase question (RAG) | 30,000 | 2,000 | $0.0015 | $1.50 |
| Long-context analysis | 150,000 | 5,000 | $0.00675 | $6.75 |
Specs
Context window
260K
tokens
Max output
66K
tokens
Input modalities
text
Tokenizer
Other