Claude Haiku 4.5 vs Nemotron 3 Ultra (free)
Side-by-side API pricing and Artificial Analysis performance for Claude Haiku 4.5 (Anthropic) and Nemotron 3 Ultra (free) (NVIDIA). Green cells mark the better value in each row.
Prices updated Jul 22, 2026, 3:55 AM UTC · refreshed hourly
| Claude Haiku 4.5 Anthropic | Nemotron 3 Ultra (free) NVIDIA | |
|---|---|---|
| Pricing | ||
| Input $/1M | $1.00 | Free |
| Output $/1M | $5.00 | Free |
| Blended $/1M (3:1) | $2.00 | Free |
| RAG example (30K in / 2K out) | $0.04 | $0 |
| Quality & speed | ||
| Intelligence Index | 29.6 | 37.8 |
| Coding Index | 43.9 | 49.3 |
| Agentic Index | 16.4 | 27.4 |
| Output speed (tok/s) | — | — |
| Time to first token | — | — |
| Specs | ||
| Context window | 200K | 1M |
| Input modalities | text, image, file | text |
FAQ
Which is cheaper, Claude Haiku 4.5 or Nemotron 3 Ultra (free)?
Nemotron 3 Ultra (free) has the lower blended API price at Free per 1M tokens (3:1 input:output mix), versus $2.00 for the other model. Prices are live from OpenRouter and refresh about hourly.
Which is better for RAG workloads?
For a typical RAG request (30K input / 2K output tokens), Nemotron 3 Ultra (free) costs about $0 per request versus $0.04. Also compare context windows: Claude Haiku 4.5 offers 200K and Nemotron 3 Ultra (free) offers 1M.
Which scores higher on Artificial Analysis benchmarks?
Nemotron 3 Ultra (free) leads on the Artificial Analysis Intelligence Index (37.8 vs 29.6). Check coding and agentic indices on this page for workload-specific tradeoffs.