AiCostCompare

Compare models

Pick up to 4 models for a side-by-side spec sheet. Hover any dotted metric name for what the test measures. Best value in each row is highlighted.

Claude Sonnet 5Gemini 3.5 Flash
Claude Sonnet 5

Anthropic

Gemini 3.5 Flash

Google

Pricing
Input $/1M$2.00$1.50
Output $/1M$10.00$9.00
Blended $/1M (3:1)$4.00$3.38
Cache read $/1M$0.20$0.15
Cache write $/1M$2.50$0.083
Quality indices
Intelligence Index38.232.6
Coding Index71.570.1
Math Index
Agentic Index43.626
Speed & latency
Output speed (tok/s)820
Time to first token118.14s0s
Time to first answer token118.14s0s
Benchmarks (Artificial Analysis)
GPQA Diamond91.1%92.2%
Humanity's Last Exam41.3%42.7%
SciCode54.3%53.9%
IFBench76.3%
AA-LCR (Long Context Reasoning)82.0%73.3%
Terminal-Bench Hard40.9%
Terminal-Bench 2.180.5%78.7%
τ²-Bench (Telecom)95.3%
τ³-Bench Banking37.3%32.2%
Design Arena (Elo)
Website1,2881,275
Web apps1,2601,219
Full-stack1,2601,208
Mobile apps1,2491,202
Android native1,2401,197
UI components1,2991,288
Data viz1,2601,247
SVG1,2121,278
3D1,2891,271
Game dev1,3131,287
Agentic game dev1,2271,180
Godot game dev1,2381,108
HTML slides1,2201,145
PPTX slides1,244
Python→PPTX slides1,2161,247
Agentic slides1,244
Agentic HTML slides1,162
Agentic slides (HTML)1,162
Agentic slides (Python PPTX)1,242
Code categories1,2931,278
ASCII art1,2261,273
Specs
Context window1M1.05M
Max output tokens128K66K
Input modalitiestext, image, filetext, image, video, file, audio
TokenizerClaudeGemini