LLM Value Ranking

Intelligence, cost and time — independently measured, weighted your way

Value ranking

Click a column header to sort
# Model Intelligence Speed tok/s Time/task (min) Tokens/task Cost/task Score
DeepSeek V4.1 Flash (Off-peak)44.97209.47.8136k¥1.65 / $0.22967.4
GPT-6.1 Sol57.9564.211.243k¥5.27 / $0.73267.0
DeepSeek V4.1 Flash (Peak)44.97209.47.8136k¥3.29 / $0.45763.0
GPT-6 Sol51.1687.17.941k¥10.75 / $1.49459.5
GPT-6 Astra60.7551.610.232k¥25.13 / $3.49152.0
GPT-6 Luna34.77128.38.566k¥0.68 / $0.09550.7
Gemini 3.8 Flash42.48248.65.485k¥15.26 / $2.11950.3
GPT-5.6 Sol49.8279.67.336k¥22.4 / $3.11150.2
GLM-5.3-Flash44.3548.920.081k¥2.02 / $0.28150.2
Mimo-V2.6 Pro47.6143.727.476k¥1.2 / $0.16747.1
GPT-5.6 Terra45.9795.39.957k¥17.63 / $2.44847.0
DeepSeek V4 Pro (Off-peak)37.29100.511.281k¥5.08 / $0.70544.5
Grok 4.643.6077.29.342k¥20.12 / $2.79443.8
Claude Opus 5.563.5092.412.7110k¥36.62 / $5.08543.7
GPT-5.6 Luna33.771248.160k¥2.33 / $0.32343.5
DeepSeek V4 Pro (Peak)37.29100.511.281k¥10.16 / $1.41140.2
GLM-5.348.7965.917.792k¥21.02 / $2.91940.1
Mimo-V2.6 Flash40.6251.230.7101k¥0.62 / $0.08638.7
Claude Sonnet 5.563.46139.112.4166k¥48.95 / $6.79936.3
Kimi K339.2733.924.756k¥14.68 / $2.03830.2
Grok 4.744.827917.680k¥38.38 / $5.3329.7
Claude Opus 553.4854.615.885k¥51.07 / $7.09329.0
Claude Fable 5.156.8468.213.085k¥56.2 / $7.80528.6
Qwen3.8 Max46.063939.5117k¥46.31 / $6.43313.1
Claude Sonnet 530.6583.819.1163k¥52.12 / $7.23910.9

Data updated 2026-10-02: intelligence, cost and time are independently measured by Artificial Analysis; prices come from official provider API pricing pages.

DeepSeek peak/off-peak share intelligence and time; only cost differs (half), so off-peak scores higher.

Prices and costs are for reference; always check official pricing.

Methodology

  1. Intelligence I: the arithmetic mean of Terminal-Bench 4.0 (agentic terminal coding), Humanity’s Last Exam (multidisciplinary reasoning) and AutomationBench (SaaS workflow automation), all independently measured by Artificial Analysis — not vendor-reported.
  2. Cost C and time T: both AA-measured: C is the full cost per benchmark task (input, cache-write, reasoning and output tokens, in USD); T is the wall-clock seconds per task. Verbose models are charged for their real token use.
  3. Composite formula: score = intelligenceNorm^a × costNorm^b × timeNorm^c (a+b+c=1). Each axis is min–max normalized to 0.05–1 across the 25 ranked models — intelligence linearly; cost and time log-scaled (floor 0.05) since they span orders of magnitude. The geometric form means a weak axis drags the total down — one extreme strength cannot buy off another axis.
  4. Presets: Balanced = ⅓/⅓/⅓; Performance = 75/12.5/12.5; Cost = 20/60/20; Speed = 20/20/60; Custom sliders combine freely (auto-normalized to 100%).
  5. Currency-independent: the score is computed from ratios, so it does not depend on currency. The cost column is measured in USD; CNY figures convert at 1 USD = 7.2 CNY for reference.
  6. Inclusion: only models with independent measurements are ranked; intelligence, speed and cost are measured values. DeepSeek off-peak shares capability and time with peak at half the cost.

FAQ

How is the composite score computed?
Score = intelligenceNorm^performance-weight × costNorm^cost-weight × timeNorm^time-weight (weights sum to 1). Each axis is first min–max normalized to 0.05–1 across the 25 ranked models — intelligence linearly; cost and time log-scaled because they span orders of magnitude — then combined geometrically: any weak axis drags the total down. All three axes are independently measured by Artificial Analysis.
Why does the ranking change completely when I switch emphasis?
Because different jobs favor different models: cheap models win one-shot tasks, strong models win multi-step critical workflows, fast models win real-time use. That is the point of this tool — find the optimum for your scenario with the weighting, not a single universal answer.
Where does the data come from?
Intelligence, full cost per task (input, cache, reasoning and output tokens) and time per task are measured by Artificial Analysis under a unified harness. Costs are in USD; CNY figures convert at 7.2 for reference.