LLM Value Ranking
Intelligence, cost and time — independently measured, weighted your way
Value ranking
| # | Model | Intelligence | Speed tok/s | Time/task (min) | Tokens/task | Cost/task | Score |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash (Off-peak) | 44.97 | 209.4 | 7.8 | 136k | ¥1.65 / $0.229 | 67.4 | |
| GPT-6.1 Sol | 57.95 | 64.2 | 11.2 | 43k | ¥5.27 / $0.732 | 67.0 | |
| DeepSeek V4.1 Flash (Peak) | 44.97 | 209.4 | 7.8 | 136k | ¥3.29 / $0.457 | 63.0 | |
| GPT-6 Sol | 51.16 | 87.1 | 7.9 | 41k | ¥10.75 / $1.494 | 59.5 | |
| GPT-6 Astra | 60.75 | 51.6 | 10.2 | 32k | ¥25.13 / $3.491 | 52.0 | |
| GPT-6 Luna | 34.77 | 128.3 | 8.5 | 66k | ¥0.68 / $0.095 | 50.7 | |
| Gemini 3.8 Flash | 42.48 | 248.6 | 5.4 | 85k | ¥15.26 / $2.119 | 50.3 | |
| GPT-5.6 Sol | 49.82 | 79.6 | 7.3 | 36k | ¥22.4 / $3.111 | 50.2 | |
| GLM-5.3-Flash | 44.35 | 48.9 | 20.0 | 81k | ¥2.02 / $0.281 | 50.2 | |
| Mimo-V2.6 Pro | 47.61 | 43.7 | 27.4 | 76k | ¥1.2 / $0.167 | 47.1 | |
| GPT-5.6 Terra | 45.97 | 95.3 | 9.9 | 57k | ¥17.63 / $2.448 | 47.0 | |
| DeepSeek V4 Pro (Off-peak) | 37.29 | 100.5 | 11.2 | 81k | ¥5.08 / $0.705 | 44.5 | |
| Grok 4.6 | 43.60 | 77.2 | 9.3 | 42k | ¥20.12 / $2.794 | 43.8 | |
| Claude Opus 5.5 | 63.50 | 92.4 | 12.7 | 110k | ¥36.62 / $5.085 | 43.7 | |
| GPT-5.6 Luna | 33.77 | 124 | 8.1 | 60k | ¥2.33 / $0.323 | 43.5 | |
| DeepSeek V4 Pro (Peak) | 37.29 | 100.5 | 11.2 | 81k | ¥10.16 / $1.411 | 40.2 | |
| GLM-5.3 | 48.79 | 65.9 | 17.7 | 92k | ¥21.02 / $2.919 | 40.1 | |
| Mimo-V2.6 Flash | 40.62 | 51.2 | 30.7 | 101k | ¥0.62 / $0.086 | 38.7 | |
| Claude Sonnet 5.5 | 63.46 | 139.1 | 12.4 | 166k | ¥48.95 / $6.799 | 36.3 | |
| Kimi K3 | 39.27 | 33.9 | 24.7 | 56k | ¥14.68 / $2.038 | 30.2 | |
| Grok 4.7 | 44.82 | 79 | 17.6 | 80k | ¥38.38 / $5.33 | 29.7 | |
| Claude Opus 5 | 53.48 | 54.6 | 15.8 | 85k | ¥51.07 / $7.093 | 29.0 | |
| Claude Fable 5.1 | 56.84 | 68.2 | 13.0 | 85k | ¥56.2 / $7.805 | 28.6 | |
| Qwen3.8 Max | 46.06 | 39 | 39.5 | 117k | ¥46.31 / $6.433 | 13.1 | |
| Claude Sonnet 5 | 30.65 | 83.8 | 19.1 | 163k | ¥52.12 / $7.239 | 10.9 |
Data updated 2026-10-02: intelligence, cost and time are independently measured by Artificial Analysis; prices come from official provider API pricing pages.
DeepSeek peak/off-peak share intelligence and time; only cost differs (half), so off-peak scores higher.
Prices and costs are for reference; always check official pricing.
Methodology
- Intelligence I: the arithmetic mean of Terminal-Bench 4.0 (agentic terminal coding), Humanity’s Last Exam (multidisciplinary reasoning) and AutomationBench (SaaS workflow automation), all independently measured by Artificial Analysis — not vendor-reported.
- Cost C and time T: both AA-measured: C is the full cost per benchmark task (input, cache-write, reasoning and output tokens, in USD); T is the wall-clock seconds per task. Verbose models are charged for their real token use.
- Composite formula: score = intelligenceNorm^a × costNorm^b × timeNorm^c (a+b+c=1). Each axis is min–max normalized to 0.05–1 across the 25 ranked models — intelligence linearly; cost and time log-scaled (floor 0.05) since they span orders of magnitude. The geometric form means a weak axis drags the total down — one extreme strength cannot buy off another axis.
- Presets: Balanced = ⅓/⅓/⅓; Performance = 75/12.5/12.5; Cost = 20/60/20; Speed = 20/20/60; Custom sliders combine freely (auto-normalized to 100%).
- Currency-independent: the score is computed from ratios, so it does not depend on currency. The cost column is measured in USD; CNY figures convert at 1 USD = 7.2 CNY for reference.
- Inclusion: only models with independent measurements are ranked; intelligence, speed and cost are measured values. DeepSeek off-peak shares capability and time with peak at half the cost.
FAQ
- How is the composite score computed?
- Score = intelligenceNorm^performance-weight × costNorm^cost-weight × timeNorm^time-weight (weights sum to 1). Each axis is first min–max normalized to 0.05–1 across the 25 ranked models — intelligence linearly; cost and time log-scaled because they span orders of magnitude — then combined geometrically: any weak axis drags the total down. All three axes are independently measured by Artificial Analysis.
- Why does the ranking change completely when I switch emphasis?
- Because different jobs favor different models: cheap models win one-shot tasks, strong models win multi-step critical workflows, fast models win real-time use. That is the point of this tool — find the optimum for your scenario with the weighting, not a single universal answer.
- Where does the data come from?
- Intelligence, full cost per task (input, cache, reasoning and output tokens) and time per task are measured by Artificial Analysis under a unified harness. Costs are in USD; CNY figures convert at 7.2 for reference.