AiCostCompare

Qwen3.8 Flash

by Alibaba (Qwen) · qwen/qwen3.8-flash · #98 cheapest of 422 paid models

Prices updated Sep 19, 2026, 6:22 PM UTC · refreshed hourly

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.15

per 1M tokens

Output

$0.47

per 1M tokens

Blended (3:1)

$0.23

3 input : 1 output

Cache read

$0.016

per 1M cached tokens

Cache write

$0.20

per 1M tokens

Cache & batch economics

Effective prices for Qwen3.8 Flash with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.23$0.0158
50%$0.18$0.00907−42%
90%$0.14$0.00371−76%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.000216$0.216
Document summary8,0001,000$0.00167$1.67
Codebase question (RAG)30,0002,000$0.00544$5.44
Long-context analysis150,0005,000$0.0249$24.85

Specs

Context window

1M

tokens

Max output

131K

tokens

Input modalities

text, image, video

Tokenizer

Qwen