GPT-3.5 Turbo (older v0613)
by OpenAI · openai/gpt-3.5-turbo-0613 · #199 cheapest of 325 paid models
Prices updated Jul 22, 2026, 1:33 AM UTC · refreshed hourly
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5Gemini 3.6 Flash
Pricing
Input
$1.00
per 1M tokens
Output
$2.00
per 1M tokens
Blended (3:1)
$1.25
3 input : 1 output
Cache read
—
per 1M cached tokens
Cache write
—
per 1M tokens
Cache & batch economics
Effective prices for GPT-3.5 Turbo (older v0613) (no cache-read rate published — hits billed as normal input).
| Cache hit rate | Effective blended $/1M | Example request* | vs no cache |
|---|---|---|---|
| 0% | $1.25 | $0.104 | — |
| 50% | $1.25 | $0.104 | −0% |
| 90% | $1.25 | $0.104 | −0% |
*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates for this provider’s batch API. Confirm on the vendor’s pricing page.
What a request costs
| Workload | Input tokens | Output tokens | Cost / request | Cost / 1K requests |
|---|---|---|---|---|
| Short chat message | 500 | 300 | $0.0011 | $1.10 |
| Document summary | 8,000 | 1,000 | $0.01 | $10.00 |
| Codebase question (RAG) | 30,000 | 2,000 | $0.034 | $34.00 |
| Long-context analysis | 150,000 | 5,000 | $0.16 | $160.00 |
Performance
Intelligence Index
—
AA composite quality
Coding Index
—
Math Index
—
Agentic Index
—
Output speed
0 tok/s
median
Time to first token
0s
median
Time to first answer
0s
after reasoning tokens
Specs
Context window
4K
tokens
Max output
4K
tokens
Input modalities
text
Tokenizer
GPT