AiCostCompare

GPT-3.5 Turbo

by OpenAI · openai/gpt-3.5-turbo · #157 cheapest of 325 paid models

Prices updated Jul 22, 2026, 1:40 AM UTC · refreshed hourly

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5Gemini 3.6 Flash

Pricing

Input

$0.50

per 1M tokens

Output

$1.50

per 1M tokens

Blended (3:1)

$0.75

3 input : 1 output

Cache read

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for GPT-3.5 Turbo (no cache-read rate published — hits billed as normal input).

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.75$0.0525
50%$0.75$0.0525−0%
90%$0.75$0.0525−0%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates for this provider’s batch API. Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.0007$0.7
Document summary8,0001,000$0.0055$5.50
Codebase question (RAG)30,0002,000$0.018$18.00
Long-context analysis150,0005,000$0.0825$82.50

Performance

Intelligence Index

3.6

AA composite quality

Coding Index

10.7

Math Index

Agentic Index

Output speed

152 tok/s

median

Time to first token

0.78s

median

Time to first answer

0.78s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
MMLU-Pro46.2%
GPQA Diamond29.7%
MATH-50044.1%

Specs

Context window

16K

tokens

Max output

4K

tokens

Input modalities

text

Tokenizer

GPT