AiCostCompare

Gemini 3.5 Flash-Lite

by Google · google/gemini-3.5-flash-lite · #165 cheapest of 325 paid models

Prices updated Jul 22, 2026, 1:31 AM UTC · refreshed hourly

Gemini 3.5 Flash-Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.30

per 1M tokens

Output

$2.50

per 1M tokens

Blended (3:1)

$0.85

3 input : 1 output

Cache read

$0.03

per 1M cached tokens

Cache write

$0.083

per 1M tokens

Cache & batch economics

Effective prices for Gemini 3.5 Flash-Lite with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.85$0.0331
50%$0.749$0.0196−41%
90%$0.668$0.0088−73%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates for this provider’s batch API. Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.0009$0.9
Document summary8,0001,000$0.0049$4.90
Codebase question (RAG)30,0002,000$0.014$14.00
Long-context analysis150,0005,000$0.0575$57.50

Performance

Intelligence Index

36.5

AA composite quality

Coding Index

49.3

Math Index

Agentic Index

26.8

Output speed

373 tok/s

median

Time to first token

5.94s

median

Time to first answer

5.94s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
GPQA Diamond83.8%
Humanity's Last Exam17.5%
SciCode40.9%
AA-LCR (Long Context Reasoning)62.0%
Terminal-Bench 2.153.6%
τ³-Bench Banking16.5%

Specs

Context window

1.05M

tokens

Max output

66K

tokens

Input modalities

text, image, video, file, audio

Tokenizer

Gemini