AiCostCompare

Gemini 3.8 Flash (batch)

by Google · google/gemini-3.8-flash:batch · #200 cheapest of 422 paid models

Prices updated Sep 20, 2026, 12:34 AM UTC · refreshed hourly

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.375

per 1M tokens

Output

$1.88

per 1M tokens

Blended (3:1)

$0.75

3 input : 1 output

Cache read

$0.037

per 1M cached tokens

Cache write

$0.042

per 1M tokens

Cache & batch economics

Effective prices for Gemini 3.8 Flash (batch) with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.75$0.0401
50%$0.623$0.0233−42%
90%$0.522$0.00975−76%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates for this provider’s batch API. Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00075$0.75
Document summary8,0001,000$0.004875$4.88
Codebase question (RAG)30,0002,000$0.015$15.00
Long-context analysis150,0005,000$0.0656$65.62

Performance

Intelligence Index

40.9

AA composite quality

Coding Index

76.3

Math Index

Agentic Index

41.1

Output speed

327 tok/s

median

Time to first token

11.27s

median

Time to first answer

11.27s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
GPQA Diamond95.3%
Humanity's Last Exam47.8%
SciCode56.6%
AA-LCR (Long Context Reasoning)81.3%
Terminal-Bench 2.187.6%
τ³-Bench Banking44.9%

Design Arena Elo

CategoryElo
Website1,315
Web apps1,256
Full-stack1,250
Mobile apps1,259
UI components1,340
Data viz1,265
3D1,323
Game dev1,339
Python→PPTX slides1,176
Code categories1,324

Specs

Context window

1.05M

tokens

Max output

66K

tokens

Input modalities

text, image, video, file, audio

Tokenizer

Gemini