AiCostCompare

Llama 4 Maverick

by Meta · meta-llama/llama-4-maverick · #90 cheapest of 325 paid models

Prices updated Jul 22, 2026, 1:31 AM UTC · refreshed hourly

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.20

per 1M tokens

Output

$0.80

per 1M tokens

Blended (3:1)

$0.35

3 input : 1 output

Cache read

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for Llama 4 Maverick (no cache-read rate published — hits billed as normal input).

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.35$0.0212
50%$0.35$0.0212−0%
90%$0.35$0.0212−0%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.00034$0.34
Document summary8,0001,000$0.0024$2.40
Codebase question (RAG)30,0002,000$0.0076$7.60
Long-context analysis150,0005,000$0.034$34.00

Performance

Intelligence Index

14.3

AA composite quality

Coding Index

16.3

Math Index

19.3

Agentic Index

1.3

Output speed

110 tok/s

median

Time to first token

0.59s

median

Time to first answer

0.59s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
MMLU-Pro80.9%
GPQA Diamond67.1%
Humanity's Last Exam4.8%
LiveCodeBench39.7%
SciCode33.1%
MATH-50088.9%
AIME39.0%
AIME 202519.3%
IFBench43.0%
AA-LCR (Long Context Reasoning)46.0%
Terminal-Bench Hard6.8%
Terminal-Bench 2.17.9%
τ²-Bench (Telecom)17.8%
τ³-Bench Banking3.9%

Design Arena Elo

CategoryElo
Website898
UI components942
Data viz919
3D958
Game dev896
Code categories913

Specs

Context window

1.05M

tokens

Max output

16K

tokens

Input modalities

text, image

Tokenizer

Llama4