AiCostCompare

DeepSeek V4 Flash 0423

by DeepSeek · deepseek/deepseek-v4-flash · #10 cheapest of 422 paid models

Prices updated Sep 20, 2026, 12:34 AM UTC · refreshed hourly

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.038

per 1M tokens

Output

$0.076

per 1M tokens

Blended (3:1)

$0.047

3 input : 1 output

Cache read

$0.0076

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for DeepSeek V4 Flash 0423 with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.047$0.003931
50%$0.036$0.002419−38%
90%$0.027$0.00121−69%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$4.16e-5$0.0416
Document summary8,0001,000$0.000378$0.378
Codebase question (RAG)30,0002,000$0.001285$1.29
Long-context analysis150,0005,000$0.006048$6.05

Performance

Intelligence Index

34.3

AA composite quality

Coding Index

69.1

Math Index

Agentic Index

27.9

Output speed

249 tok/s

median

Time to first token

0.93s

median

Time to first answer

8.96s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
GPQA Diamond90.8%
Humanity's Last Exam38.6%
SciCode50.3%
AA-LCR (Long Context Reasoning)79.7%
Terminal-Bench 2.178.7%
τ³-Bench Banking39.4%

Design Arena Elo

CategoryElo
Website1,220
UI components1,179
Data viz1,142
SVG1,179
3D1,217
Game dev1,220
Code categories1,220
ASCII art1,130

Specs

Context window

1.05M

tokens

Max output

384K

tokens

Input modalities

text

Tokenizer

DeepSeek