HiseHub Model Value Lab
AI model comparison helps answer which model is actually worth using. Compare intelligence, agentic and coding signals, output speed, task cost, and token pricing in one place, then connect the result to real subscription choices.
HiseHub Model Value Lab
Compare AI models by intelligence, speed, and API price
Use independent benchmark data to understand which models are stronger, faster, or more economical for different kinds of work. This is a decision aid—not a promise that one model wins every task.
Side-by-side comparison
Compare up to three models
Select models from the table to compare the cached Artificial Analysis fields on the same snapshot. Search and sorting do not clear your choices.
| Compare | Rank | Model | Creator | Intelligence | Agentic | Coding | Cost / task | Output speed | Input / output |
|---|---|---|---|---|---|---|---|---|---|
| #1 | Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Released 2026-09-01 | Anthropic | 53.4 | 58.0 | 81.6 | $7.630 | 69.4 tok/sTTFT 275.52s · E2E 282.73s | $10.00 / $50.00 per 1MCache $0.25 / 1M | |
| #2 | GPT-6 Astra (max)Released 2026-09-03 | OpenAI | 52.8 | 51.5 | 76.9 | $3.258 | 56.5 tok/sTTFT 322.48s · E2E 331.33s | $10.00 / $50.00 per 1MCache $1.00 / 1M | |
| #3 | GPT-6 Astra (xhigh)Released 2026-09-03 | OpenAI | 52.5 | 50.6 | 75.9 | $2.309 | 52.4 tok/sTTFT 168.66s · E2E 178.21s | $10.00 / $50.00 per 1MCache $1.00 / 1M | |
| #4 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Released 2026-07-24 | Anthropic | 49.7 | 55.5 | 77.0 | $4.878 | 52.0 tok/sTTFT 28.65s · E2E 38.27s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #5 | Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Released 2026-09-01 | Anthropic | 49.1 | 50.6 | 77.1 | $2.983 | 55.3 tok/sTTFT 7.99s · E2E 17.03s | $10.00 / $50.00 per 1MCache $0.25 / 1M | |
| #6 | Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Released 2026-09-01 | Anthropic | 47.0 | 47.8 | 75.2 | $2.371 | 51.8 tok/sTTFT 6.17s · E2E 15.83s | $10.00 / $50.00 per 1MCache $0.25 / 1M | |
| #7 | GPT-6 Astra (Non-reasoning)Released 2026-09-03 | OpenAI | 45.2 | 47.8 | 76.2 | $1.712 | Not provided | $10.00 / $50.00 per 1MCache $1.00 / 1M | |
| #8 | Grok 4.6 (medium)Released 2026-08-12 | SpaceXAI | 43.0 | 51.2 | 74.4 | $1.496 | 60.5 tok/sTTFT 31.43s · E2E 39.69s | $2.00 / $6.00 per 1MCache $0.50 / 1M | |
| #9 | GLM-5.3-FlashReleased 2026-08-26 | Z AI | 41.9 | 51.2 | 71.5 | $0.253 | 66.5 tok/sTTFT 1.89s · E2E 39.52s | $0.15 / $0.50 per 1MCache $0.03 / 1M | |
| #10 | Muse Spark 1.2 (xhigh)Released 2026-08-05 | Meta | 39.8 | 44.0 | 72.2 | $0.975 | 237.9 tok/sTTFT 14.05s · E2E 24.55s | $1.25 / $4.25 per 1MCache $0.15 / 1M | |
| #11 | Claude Opus 5 (Adaptive Reasoning, Low Effort)Released 2026-07-24 | Anthropic | 39.8 | 37.5 | 66.9 | $1.098 | 53.0 tok/sTTFT 2.37s · E2E 11.80s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #12 | GPT-5.5 (xhigh)Released 2026-04-23 | OpenAI | 38.6 | 37.3 | 74.9 | $2.634 | 90.4 tok/sTTFT 58.94s · E2E 64.47s | $5.00 / $30.00 per 1MCache $0.50 / 1M | |
| #13 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Released 2026-06-30 | Anthropic | 38.4 | 44.3 | 71.5 | $5.091 | 79.5 tok/sTTFT 178.48s · E2E 184.77s | $2.00 / $10.00 per 1MCache $0.20 / 1M | |
| #14 | GPT-5.6 Luna (max)Released 2026-07-09 | OpenAI | 37.5 | 42.7 | 71.4 | $0.178 | 119.8 tok/sTTFT 168.22s · E2E 172.39s | $0.20 / $1.20 per 1MCache $0.02 / 1M | |
| #15 | GPT-5.6 Sol (low)Released 2026-07-09 | OpenAI | 33.8 | 32.6 | 69.7 | $0.261 | 67.4 tok/sTTFT 2.71s · E2E 10.13s | $4.00 / $20.00 per 1MCache $0.40 / 1M | |
| #16 | Gemini 3.5 Flash (high)Released 2026-05-19 | 33.0 | 27.3 | 70.1 | $1.563 | 216.6 tok/sTTFT 16.09s · E2E 18.40s | $1.50 / $9.00 per 1MCache $0.15 / 1M | ||
| #17 | GPT-5.6 Terra (medium)Released 2026-07-09 | OpenAI | 32.8 | 31.4 | 64.7 | Not provided | 93.8 tok/sTTFT 1.63s · E2E 6.96s | $2.00 / $12.00 per 1MCache $0.20 / 1M | |
| #18 | Kimi K2.6Released 2026-04-20 | Kimi | 31.3 | 22.1 | 61.8 | Not provided | 56.5 tok/sTTFT 2.78s · E2E 90.49s | $0.95 / $4.00 per 1MCache $0.16 / 1M | |
| #19 | Claude Opus 4.7 (Non-reasoning, High Effort)Released 2026-04-16 | Anthropic | 30.9 | Not provided | Not provided | Not provided | 46.5 tok/sTTFT 1.20s · E2E 11.96s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #20 | Apodex 1.1Released 2026-08-30 | Apodex | 30.4 | Not provided | 60.8 | Not provided | Not provided | $0.30 / $3.00 per 1MCache $0.03 / 1M | |
| #21 | GPT-5.2 (xhigh)Released 2025-12-11 | OpenAI | 30.4 | Not provided | Not provided | Not provided | 88.9 tok/sTTFT 121.46s · E2E 127.08s | $1.75 / $14.00 per 1MCache $0.17 / 1M | |
| #22 | MiniMax-M3Released 2026-06-01 | MiniMax | 29.6 | 30.8 | 58.6 | $0.508 | 155.5 tok/sTTFT 1.31s · E2E 17.39s | $0.30 / $1.20 per 1MCache $0.06 / 1M | |
| #23 | Claude Opus 4.5 (Reasoning)Released 2025-11-24 | Anthropic | 29.1 | Not provided | Not provided | Not provided | 51.9 tok/sTTFT 14.90s · E2E 24.52s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #24 | GPT-5.2 Codex (xhigh)Released 2025-12-11 | OpenAI | 28.5 | Not provided | Not provided | Not provided | Not provided | $1.75 / $14.00 per 1MCache $0.17 / 1M | |
| #25 | Solar Pro 4Released 2026-08-06 | Upstage | 28.2 | Not provided | 52.7 | Not provided | 80.3 tok/sTTFT 1.98s · E2E 33.12s | $0.30 / $1.20 per 1MCache $0.06 / 1M | |
| #26 | Nex-N2-ProReleased 2026-06-02 | Nex AGI | 28.2 | Not provided | 59.1 | Not provided | 115.3 tok/sTTFT 1.75s · E2E 23.44s | $0.50 / $2.50 per 1MCache $0.25 / 1M | |
| #27 | GLM-5 (Reasoning)Released 2026-02-11 | Z AI | 27.9 | Not provided | Not provided | Not provided | 71.3 tok/sTTFT 1.42s · E2E 51.97s | $1.00 / $3.20 per 1MCache $0.20 / 1M | |
| #28 | Grok Build 0.1 0616Released 2026-06-16 | SpaceXAI | 27.2 | Not provided | 51.5 | Not provided | 73.1 tok/sTTFT 0.50s · E2E 34.69s | $1.00 / $2.00 per 1MCache $0.20 / 1M | |
| #29 | GLM-5-TurboReleased 2026-03-15 | Z AI | 26.6 | Not provided | Not provided | Not provided | Not provided | Not provided | |
| #30 | MiMo-V2.5-ProReleased 2026-04-22 | Xiaomi | 26.4 | 22.7 | 60.2 | $0.054 | 36.2 tok/sTTFT 3.55s · E2E 72.58s | $0.43 / $0.87 per 1MCache $0.00 / 1M | |
| #31 | Claude Opus 4.6 (Non-reasoning, High Effort)Released 2026-02-05 | Anthropic | 26.4 | Not provided | Not provided | Not provided | 39.0 tok/sTTFT 2.05s · E2E 14.87s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #32 | Inkling (xhigh)Released 2026-07-15 | Thinking Machines | 25.5 | 24.3 | 52.1 | $0.607 | 71.5 tok/sTTFT 2.51s · E2E 37.45s | $1.00 / $4.05 per 1MCache $0.17 / 1M | |
| #33 | Grok 4.20 0309 (Reasoning)Released 2026-03-10 | SpaceXAI | 25.2 | Not provided | Not provided | Not provided | Not provided | $2.00 / $6.00 per 1MCache $0.20 / 1M | |
| #34 | MiMo-V2-Omni-0327Released 2026-03-27 | Xiaomi | 25.1 | Not provided | Not provided | Not provided | Not provided | Not provided | |
| #35 | Claude Sonnet 5 (Adaptive Reasoning, Low Effort)Released 2026-06-30 | Anthropic | 24.7 | Not provided | Not provided | $0.509 | 62.0 tok/sTTFT 1.33s · E2E 9.40s | $2.00 / $10.00 per 1MCache $0.20 / 1M | |
| #36 | Claude Sonnet 4.6 (Non-reasoning, High Effort)Released 2026-02-17 | Anthropic | 24.7 | Not provided | Not provided | Not provided | 44.3 tok/sTTFT 1.81s · E2E 13.10s | $3.00 / $15.00 per 1MCache $0.30 / 1M | |
| #37 | Gemini 3.5 Flash (minimal)Released 2026-05-19 | 23.8 | Not provided | Not provided | Not provided | 196.6 tok/sTTFT 1.00s · E2E 3.54s | $1.50 / $9.00 per 1MCache $0.15 / 1M | ||
| #38 | GPT-5.1 Codex (high)Released 2025-11-13 | OpenAI | 23.7 | Not provided | Not provided | Not provided | Not provided | $1.25 / $10.00 per 1M | |
| #39 | Claude Opus 4.5 (Non-reasoning)Released 2025-11-24 | Anthropic | 23.7 | Not provided | Not provided | Not provided | 48.1 tok/sTTFT 1.34s · E2E 11.74s | $5.00 / $25.00 per 1MCache $0.50 / 1M | |
| #40 | Kimi K2.6 (Non-reasoning)Released 2026-04-20 | Kimi | 23.6 | Not provided | Not provided | Not provided | 51.4 tok/sTTFT 2.62s · E2E 12.35s | $0.95 / $4.00 per 1MCache $0.16 / 1M | |
| #41 | GLM 5V Turbo (Reasoning)Released 2026-04-01 | Z AI | 23.5 | Not provided | Not provided | Not provided | Not provided | Not provided | |
| #42 | GPT-5 (high)Released 2025-08-07 | OpenAI | 23.0 | Not provided | 37.8 | Not provided | 100.8 tok/sTTFT 68.57s · E2E 73.53s | $1.25 / $10.00 per 1MCache $0.13 / 1M | |
| #43 | Qwen3.5 27B (Reasoning)Released 2026-02-24 | Alibaba | 22.9 | Not provided | Not provided | Not provided | 77.3 tok/sTTFT 5.62s · E2E 37.96s | $0.30 / $2.40 per 1M | |
| #44 | MiniMax-M2.5Released 2026-02-12 | MiniMax | 22.8 | Not provided | Not provided | Not provided | 96.2 tok/sTTFT 1.63s · E2E 27.62s | $0.30 / $1.20 per 1MCache $0.03 / 1M | |
| #45 | Gemini 3.5 Flash-LiteReleased 2026-07-21 | 22.7 | 15.9 | 49.3 | $0.124 | 362.7 tok/sTTFT 7.44s · E2E 8.82s | $0.30 / $2.50 per 1MCache $0.03 / 1M | ||
| #46 | MiMo-V2-Flash (Feb 2026)Released 2025-12-16 | Xiaomi | 22.4 | Not provided | Not provided | Not provided | Not provided | Not provided | |
| #47 | Qwen3.8 27B (Non-reasoning)Released 2026-08-14 | Alibaba | 22.4 | Not provided | 44.6 | Not provided | 55.8 tok/sTTFT 3.83s · E2E 12.79s | $0.50 / $3.00 per 1MCache $0.05 / 1M | |
| #48 | MiMo-V2.5Released 2026-04-22 | Xiaomi | 22.3 | 17.4 | 56.8 | $0.019 | 65.5 tok/sTTFT 6.11s · E2E 44.31s | $0.14 / $0.28 per 1MCache $0.00 / 1M | |
| #49 | GPT-5.6 Luna (low)Released 2026-07-09 | OpenAI | 21.8 | 17.9 | 44.2 | Not provided | 115.8 tok/sTTFT 1.79s · E2E 6.11s | $0.20 / $1.20 per 1MCache $0.02 / 1M | |
| #50 | Qwen3.5 397B A17B (Non-reasoning)Released 2026-02-16 | Alibaba | 21.4 | Not provided | Not provided | Not provided | 82.6 tok/sTTFT 2.12s · E2E 8.17s | $0.60 / $3.60 per 1M |
Last synchronized: . Rankings and prices can change as providers and benchmark versions change.
How to read this comparison
- Intelligence is the Artificial Analysis Intelligence Index score; it is a composite benchmark, not a universal quality guarantee.
- Value signal is a simple HiseHub comparison of intelligence against the estimated 3:1 input/output price. It is not a provider quote or a recommendation to buy API credits.
- Speed and latency are median measurements. Real response time also depends on load, routing, prompt length, and the task.
- Coverage includes the public language-model fields available through the official free API. Some models do not yet have every benchmark score, so a dash means “not measured” rather than zero.
- Consumer subscription prices and regional App Store offers are separate from API prices. See our AI subscription price tool before choosing a plan.
Data source: Artificial Analysis. HiseHub does not reproduce its full benchmark suite and is not affiliated with Artificial Analysis or any model provider.
Compare models before choosing a plan
A leaderboard is useful when it explains what the number measures. HiseHub separates model capability from consumer subscription access: a high benchmark score does not automatically mean that every ChatGPT, Claude, Gemini, or Grok plan includes that model or the same usage limits.
What the comparison measures
- Intelligence: a composite score from independent Artificial Analysis evaluations. It is useful for a broad shortlist, not a universal winner.
- Agentic and coding: capability signals derived from the published evaluation data when available. A dash means that a metric was not measured for that model.
- Speed and latency: median output tokens per second and related response measurements. Real-world response time also depends on routing, load, prompt length, and task design.
- Cost per task: the source’s estimated cost for its Intelligence Index task mix, which is more informative than token price alone but still not a personal usage quote.
- Input and output price: API token prices per million tokens. The value sort uses a clearly labelled HiseHub 3:1 input/output estimate when both prices are available.
- Release date: helps distinguish current model versions from older variants with similar names.
Use the right model for the job
Use the search box to find a model or creator, then sort the table for the question you actually have. For writing and editing, start with intelligence and instruction-following evidence; for coding, inspect coding and agentic signals; for research, check reliability, long context, and citation behaviour; for images and video, use a media-specific evaluation instead of assuming a language-model score transfers.
| Task | What to compare first | What to verify before paying |
|---|---|---|
| Writing and editing | Instruction following, quality, and speed | Whether the consumer plan includes the model and enough usage |
| Coding | Coding score, tool use, context, and reliability | Whether the product or API route is billed separately |
| Research | Knowledge reliability, long context, and citations | Search, file, and deep-research limits on your plan |
| Images and video | Blind preference scores and generation limits | Credits, regional availability, and commercial-use terms |
Why rankings can disagree
Different evaluations reward different behaviour. A model that is excellent at coding may be less convenient for everyday writing; a fast model may trade depth for latency; and a cheaper API can consume more tokens to complete the same task. Treat the table as a starting point, not as a universal winner. When two models are close, your own representative prompts are often the best tie-breaker.
How HiseHub keeps the data useful
The tool uses the official Artificial Analysis public language-model API, fetches all available pages within the free endpoint, stores the result on the server, and shows the refresh time. The API key is never sent to site visitors. The number of rows and the Intelligence Index version can change as the source adds, retires, or re-evaluates models. Rows without a particular benchmark remain visible so the table does not silently hide the wider model catalogue.
Artificial Analysis requires visible source attribution and its terms govern API use. HiseHub is an independent comparison layer; it does not reproduce the source site’s full benchmark suite or imply endorsement. Read the Artificial Analysis methodology when you need the definitions behind a metric.
For consumer subscription pricing by country, use the AI subscription price comparison. For caps, credits, and reset windows, browse the AI usage limits hub.
Questions readers ask
Does the highest score mean the best model?
No. It means the model performed strongly on the selected evaluation mix. Your workload, preferred interface, context needs, privacy requirements, and budget still matter.
Is API price the same as a ChatGPT or Claude subscription?
No. API pricing is usage-based and usually quoted per million tokens. A consumer subscription may bundle a model with a monthly fee, message caps, file limits, or temporary capacity. Compare the two routes separately before purchasing.
Why does the same model look different across providers?
Hosted endpoints can use different quantization, routing, sampling defaults, context settings, and capacity. Provider-level performance is therefore a separate question from the model name alone.
Use the ranking as a shortlist
The practical workflow is simple: shortlist two or three models, test them on the same representative prompt, check the official plan or API limits, and then choose the cheapest route that meets your quality and reliability needs. Recheck the data after a major model release or pricing change.
HiseHub is independent. Benchmark results and API prices are point-in-time information and may change. Always confirm the current model, plan, price, and usage limits on the provider’s official page.
How to interpret this AI model comparison
This table is a decision aid, not a universal winner list. Intelligence, agentic and coding signals come from the published benchmark snapshot; output speed and API price describe different trade-offs. Compare models within the same task and date before drawing a conclusion.
For a buying decision, first define the work you need to do, then check the model’s access route, usage limits and price. A model can rank highly while being unavailable on the plan or in the region you use. Recheck the linked source after a major model release or pricing change.
New model announced but no independent score yet?
Use the AI Model Release Tracker to compare official release dates, prices, context and rollout status before a matching benchmark row arrives.
Considering a self-hosted model stack?
After comparing model capability and price, use the Self-Hosted AI Tools directory to compare local runtimes, chat interfaces, and workflow builders.