AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

R1 Distill Llama 70BGPT-5.6 Sol
R1 Distill Llama 70B

DeepSeek

GPT-5.6 Sol

OpenAI

Pricing
Input $/1M$0.80$5.00
Output $/1M$0.80$30.00
Blended $/1M (3:1)$0.80$11.25
Cache read $/1M$0.50
Cache write $/1M$6.25
Quality indices
Intelligence Index9.958.9
Coding Index77.4
Math Index53.7
Agentic Index54
Speed & latency
Output speed (tok/s)3267
Time to first token0.61s120.38s
Time to first answer token62.16s120.38s
Benchmarks (Artificial Analysis)
MMLU-Pro79.5%
GPQA Diamond40.2%94.1%
Humanity's Last Exam6.1%47.2%
LiveCodeBench26.6%
SciCode31.3%56.1%
MATH-50093.5%
AIME67.0%
AIME 202553.7%
IFBench27.6%72.7%
AA-LCR (Long Context Reasoning)11.0%73.7%
Terminal-Bench Hard1.5%65.9%
Terminal-Bench 2.188.0%
τ²-Bench (Telecom)21.9%85.1%
τ³-Bench Banking33.0%
Specs
Context window8K1.05M
Max output tokens8K128K
Input modalitiestextfile, image, text
TokenizerLlama3GPT