AiCostCompare

Llama 4 Maverick

by Meta · meta-llama/llama-4-maverick · #115 cheapest of 422 paid models

Prices updated Sep 19, 2026, 6:21 PM UTC · refreshed hourly

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Compare vs:Claude Sonnet 5Claude Opus 4.6Claude Haiku 4.5GPT-5.6 Sol

Pricing

Input

$0.188

per 1M tokens

Output

$0.652

per 1M tokens

Blended (3:1)

$0.304

3 input : 1 output

Cache read

per 1M cached tokens

Cache write

per 1M tokens

Cache & batch economics

Effective prices for Llama 4 Maverick (no cache-read rate published — hits billed as normal input).

Cache hit rateEffective blended $/1MExample request*vs no cache
0%$0.304$0.0198
50%$0.304$0.0198−0%
90%$0.304$0.0198−0%

*Example: 100K sticky context + 2K new input + 1K output. Batch discount is an approximate 50% off list rates (available for OpenAI / Anthropic / Google — not flagged for this creator). Confirm on the vendor’s pricing page.

What a request costs

WorkloadInput tokensOutput tokensCost / requestCost / 1K requests
Short chat message500300$0.000289$0.2895
Document summary8,0001,000$0.002152$2.15
Codebase question (RAG)30,0002,000$0.00693$6.93
Long-context analysis150,0005,000$0.0314$31.39

Performance

Intelligence Index

9.3

AA composite quality

Coding Index

16.3

Math Index

19.3

Agentic Index

0.6

Output speed

111 tok/s

median

Time to first token

0.57s

median

Time to first answer

0.57s

after reasoning tokens

Artificial Analysis benchmarks

BenchmarkScore
MMLU-Pro80.9%
GPQA Diamond67.1%
Humanity's Last Exam4.9%
LiveCodeBench39.7%
SciCode31.7%
MATH-50088.9%
AIME39.0%
AIME 202519.3%
IFBench43.0%
AA-LCR (Long Context Reasoning)50.0%
Terminal-Bench Hard6.8%
Terminal-Bench 2.17.9%
τ²-Bench (Telecom)17.8%
τ³-Bench Banking3.7%

Design Arena Elo

CategoryElo
Website883
UI components914
Data viz895
3D929
Game dev861
Code categories895

Specs

Context window

1.05M

tokens

Max output

16K

tokens

Input modalities

text, image

Tokenizer

Llama4