AiCostCompare

GPT-4.1 Nano

by OpenAI · openai/gpt-4.1-nano · #54 cheapest of 325 paid models

Prices updated Jul 22, 2026, 1:39 AM UTC · refreshed hourly

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5Gemini 3.6 Flash

Pricing

Input

$0.10

per 1M tokens

Output

$0.40

per 1M tokens

Blended (3:1)

$0.175

3 input : 1 output

Cache read

$0.025

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for GPT-4.1 Nano with prompt caching.

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.175$0.0106
50%$0.147$0.00685−35%
90%$0.124$0.00385−64%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates for this provider’s batch API. Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00017$0.17
Document summary8,0001,000$0.0012$1.20
Codebase question (RAG)30,0002,000$0.0038$3.80
Long-context analysis150,0005,000$0.017$17.00

Performance

Intelligence Index

9.6

AA composite quality

Coding Index

11.1

Math Index

24

Agentic Index

1.2

Output speed

189 tok/s

median

Time to first token

0.42s

median

Time to first answer

0.42s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
MMLU-Pro65.7%
GPQA Diamond51.2%
Humanity's Last Exam3.9%
LiveCodeBench32.6%
SciCode25.9%
MATH-50084.8%
AIME23.7%
AIME 202524.0%
IFBench32.0%
AA-LCR (Long Context Reasoning)17.0%
Terminal-Bench Hard3.8%
Terminal-Bench 2.13.7%
τ²-Bench (Telecom)17.3%
τ³-Bench Banking3.5%

Design Arena Elo

CategoryElo
Website1,000
UI components959
Data viz928
3D983
Game dev1,029
Code categories996

Specs

Context window

1.05M

tokens

Max output

33K

tokens

Input modalities

image, text, file

Tokenizer

GPT