Rankings · Updated July 2026

AI model rankings

Compare 38 frontier, mid-tier, and open-weight models. Start with capability, or switch to Value to include price. Capability profiles are editorial judgments, not independent measurements. Speed stays blank unless a reproducible or clearly attributed figure exists.

Data from models.json Rankings computed from data — never paid placements
Type
License
Rank by

# Model benchr Rating SWE-bench % Input $/1M Output $/1M Context Tok/s Released
1Claude Fable 5Anthropicfrontier9.8—$10.00$50.001M—Jun 2026
2Claude Fable 5.1Anthropicfrontier9.8—$10.00$50.001M—Sep 2026
3Claude Opus 5Anthropicfrontier9.6—$5.00$25.001M—Jul 2026
4Claude Opus 4.7Anthropicfrontier9.587.6%$5.00$25.001M—Apr 2026
5Claude Opus 4.8Anthropicfrontier9.588.6%$5.00$25.001M—May 2026
6Claude Sonnet 5Anthropicfrontier9.589.4%$2.00$10.001M—Jul 2026
7GPT-5.6OpenAIfrontier9.489.8%$4.00$20.001.1M—Jul 2026
8GPT-5.5OpenAIfrontier9.3—$5.00$30.001.1M—Apr 2026
9Kimi K3Moonshot AIfrontier open9.2—$3.00$15.001.0M—Jul 2026
10Grok 4.6xAIfrontier9.2—$2.00$6.00500K—Aug 2026
11DeepSeek V4-ProDeepSeekfrontier open9.180.6%$1.32$3.961M—Apr 2026
12GLM-5.3Z.AIfrontier9.1—$1.40$4.401M—Aug 2026
13Grok 4.5xAIfrontier9.1—$2.00$6.00500K—Jul 2026
14GPT-5.4OpenAIfrontier9.0—$2.50$15.001.1M—Mar 2026
15Gemini 3.7 FlashGooglemid8.9—$0.750$3.751.0M—Aug 2026
16Claude Sonnet 4.6Anthropicmid8.879.6%$3.00$15.001M—Feb 2026
17GPT-5OpenAIfrontier8.874.9%$1.25$10.00400K—Aug 2025
18GLM-5.2Z.AIfrontier open8.8—$1.40$4.401M—Jun 2026
19Gemini 3.6 FlashGooglemid8.7—$0.750$3.751.0M—Jul 2026
20MiniMax M3MiniMaxfrontier open8.7—$0.300$1.201M—Jun 2026
21Qwen3.6-27BAlibaba (Qwen)open8.677.2%FreeFree262K—Apr 2026
22Gemini 3.1 ProGooglefrontier8.680.6%$2.00$12.001M—Feb 2026
23Gemini 3.5 FlashGooglemid8.6—$1.50$9.001.0M—May 2026
24Mistral Medium 3.5Mistral AIfrontier open8.6—$1.50$7.50256K—Apr 2026
25Gemini 3.5 Flash-LiteGooglefrontier8.2—$0.300$2.501.0M—Jul 2026
26Grok 4.3xAIfrontier8.2—$1.25$2.501M——
27Llama 4 MaverickMetafrontier open8.0—FreeFree1M—Apr 2025
28Mistral Large 3Mistral AIfrontier open7.8—$0.500$1.50256K—Dec 2025
29Kimi K2.6Moonshot AIfrontier open7.880.2%$0.950$4.00262K——
30Claude Haiku 4.5Anthropicsmall7.673.3%$1.00$5.00200K—Oct 2025
31Phi-4Microsoftsmall open7.4—FreeFree16K—Dec 2024
32Llama 4 ScoutMetaopen7.3—FreeFree10M—Apr 2025
33GPT-5 MiniOpenAIsmall7.3—$0.250$2.00400K—Aug 2025
34DeepSeek-V4.1-FlashDeepSeekopen0.0—$0.300$1.201M—Sep 2026
35Gemini 3.8 FlashGooglemid0.0—$0.750$3.751.0M—Sep 2026
36GPT-6 AstraOpenAIfrontier0.0—$10.00$50.001.1M—Sep 2026
37Qwen3.8-MaxAlibaba (Qwen)frontier0.0—$2.00$6.001M—Aug 2026
38Qwen3.8-FlashAlibaba (Qwen)mid0.0—$0.150$0.470262K—Aug 2026

How the benchr Rating works

The Rank by toggle gives the benchr Rating two meanings. Quality uses capability alone. Value also includes listed API rates. Neither is a lab measurement or a reader poll; both are transparent editorial calculations, and no provider pays for placement.

Quality — the default

This view combines coding, reasoning, and writing. Its inputs come from the public record and are inspectable in models.json.

capability = (coding × 0.40) + (reasoning × 0.40) + (writing × 0.20)

Value — the optional lens

This view blends capability with API-rate efficiency. Price is the average of listed input and output rates per million tokens. A self-hosted model with no per-token API rate receives the maximum API-price score. Hardware, operations, and hosted inference are excluded, so Value is not a total-cost comparison.

blended = (input_per_million + output_per_million) / 2 price_score = max(0, min(100, 100 × (1 − max(0, blended − 0.50) / 29.50))) value_score = round(capability × 0.65 + price_score × 0.35)

Both run in assets/js/models.js, so you can read and verify them yourself. Each produces a 0–100 value shown on a 0–10 scale, and price is shown directly in the Input/Output columns regardless of mode. For verified official pricing and benchmark figures, see model-figures.json.

Methodology as of June 1, 2026. The formula may change as the model field changes; check the changelog for updates.

Frequently asked questions

What is the benchr Rating?

By default it's pure capability (coding, reasoning, writing). A “Rank by” toggle switches it to a Value lens that also folds in price efficiency (most capability per dollar). Both run in open JavaScript you can read in assets/js/models.js; price is shown in its own columns too.

Which AI model is best in 2026?

It depends on the task and budget. The default benchr Rating is an editorial capability lens, while Value blends that lens with listed prices. Use the rankings to form a shortlist and validate it on representative tasks; no score on this page establishes a universal winner.

Are rankings ever paid or sponsored?

No. Rankings are computed purely from data in models.json. No provider has paid for placement. See editorial standards for the full policy.

How often is the data updated?

The updated field in models.json shows the last data refresh. Major releases and price changes are added after they are checked against provider documentation. Spot an error? File a correction.

What is the difference between a benchmark result and a capability profile?

A published benchmark result names the evaluation variant and its source. The 0–100 capability profiles are benchr editorial judgments informed by cited public evidence; they are not benchmark results measured or independently reproduced by benchr. Missing official values stay blank. The methodology page explains the distinction, and provider facts are tracked in model-figures.json.

Other tools

Charts → Intelligence-vs-price quadrant and weighted benchmark explorer Cost calculator → Enter your token usage, get monthly cost ranked cheapest-first Model recommender → Answer three questions, get your best-fit pick with a reason Side-by-side compare → Pick up to five models and compare every dimension Pricing Index → View complete model pricing and cost optimization guide Cheapest Leaderboard → Compare cheapest AI models by token cost Coding Leaderboard → Rank coding models by SWE-bench Verified Reasoning Leaderboard → Rank models by GPQA and reasoning capabilities Context Window Leaderboard → Rank models by context token sizes
Updates

Follow new pieces through RSS, recent releases, and the changelog.