What this guide covers
This guide covers the hosted frontier candidates as of September 2026: Anthropic's Claude Opus 5 and Claude Fable 5.1, OpenAI's GPT-5.6 Sol and GPT-6 Astra, Google's Gemini 3.1 Pro (Google still labels it Preview), and xAI's Grok 4.6, with DeepSeek-V4-Pro and Kimi K3 as open-weight checks. The reviews and the comparison below were written for the earlier flagships — Claude Opus 4.7 and GPT-5 — and stay as the record of those models; their evaluation method carries over unchanged. The linked pages separate provider-published specifications and benchmarks from editorial hypotheses, and none claims an unpublished benchr test or a universal winner.
The current shortlist
| Model | Price / 1M | Context | What the provider publishes |
|---|---|---|---|
| Claude Opus 5 | $5 / $25 | 1M | Anthropic's starting point for most workloads; benchr records no Anthropic benchmark table for it. |
| Claude Fable 5.1 | $10 / $50 | 1M | Anthropic points demanding reasoning and long-horizon agentic work here. The API returns HTTP 400 when tool_choice is any or a named tool. |
| GPT-6 Astra | $10 / $50 | 1.05M | OpenAI's most capable model; no benchmark table published. Tool calling requires the Responses API. |
| GPT-5.6 Sol | $4 / $20 | 1.05M | SWE-bench Verified 89.8, GPQA Diamond 91.2 (OpenAI-reported). |
| Gemini 3.1 Pro | $2 / $12 up to 200K-token prompts; $4 / $18 above | ~1M | GPQA Diamond 94.3, SWE-bench Verified 80.6 (Google-reported). Still Google's newest Pro model. |
| Grok 4.6 | $2 / $6 under 200K-token prompts; $4 / $12 at or above | 500K | xAI's flagship; benchr records no provider benchmark figure for it. |
| DeepSeek-V4-Pro (open weights, MIT) | $1.32 / $3.96 peak; $0.66 / $1.98 off-peak | 1M | SWE-bench Verified 80.6, Terminal-Bench 2.1 87.9 (DeepSeek-reported). |
| Kimi K3 (Kimi K3 License) | $3 / $15 | 1M | Terminal-Bench 2.1 88.3, GPQA 93.5 (Moonshot-reported); no SWE-bench Verified figure. |
Reviews
-
Claude Opus 4.7, reviewed
Anthropic positions Opus 4.7 for demanding coding and agentic work and publishes its price and benchmark sheet. Treat architectural fit as an editorial hypothesis: test it on held-out repository issues and document tasks, then compare errors and review time.
-
GPT-5, reviewed
OpenAI publishes GPT-5 specifications and benchmark results that make it a candidate for structured output and reasoning workloads. Speed, prose preference, and technical error rate depend on the deployment and prompt; measure them on the same local rubric.
-
Gemini 3.1 Pro, reviewed
Google documents the model's context capacity and Workspace integrations. Image-heavy and Workspace-bound work are reasonable editorial hypotheses to test with representative screenshots, documents, citations, and failure cases.
Comparisons
-
GPT-5 vs Claude Opus 4.7: a seven-task evaluation plan
Seven workload categories with a shared prompt and scoring template. The category leans are editorial hypotheses, not private measured results; reproduce them with pinned versions, retained outputs, and blind review. Written for the earlier pair, the plan applies unchanged to Claude Opus 5 and GPT-5.6 Sol.
-
Multimodal evaluation plan: twelve images, four models
A reproducible rubric for Claude, GPT-5, Gemini 3, and Llama 4 across dense interfaces, document images, charts, and Arabic script. Placements are hypotheses to validate, not an unpublished winner table.
-
The price-per-use-case table
Six workload shapes using dated, provider-listed token prices. Recalculate with your input/output mix, cache hit rate, tool calls, retries, and current price pages before making a budget decision.
Which one should you use?
Start with the candidate whose provider-documented capabilities match your dominant workload. For coding and agentic work, put GPT-5.6 Sol (SWE-bench Verified 89.8, OpenAI-reported) next to Claude Opus 5, which Anthropic names as the model to start with for most workloads, though benchr records no Anthropic benchmark table for it. For multimodal, long-document, or Workspace-bound work, test Gemini 3.1 Pro (GPQA Diamond 94.3, Google-reported). Add Grok 4.6 when output cost dominates: at $2 / $6 per 1M tokens under 200K-token prompts, it has the lowest output rate among the closed models on this shortlist. These are editorial starting hypotheses, not measured winners.
Reserve GPT-6 Astra and Claude Fable 5.1, both $10 / $50 per 1M tokens, for tasks where your own evaluation shows the ceiling pays for itself. Check the integration first: GPT-6 Astra requires the Responses API for tool calling and accepts no custom temperature, and Claude Fable 5.1 returns HTTP 400 if tool_choice is any or a named tool. If the weights must stay on your own hardware, use DeepSeek-V4-Pro (MIT license, SWE-bench Verified 80.6, DeepSeek-reported) as the open-weight control.
To decide whether two subscriptions add value, route the same held-out set through each model alone and through the proposed pair. Record incremental task coverage, review time, failures, and the actual monthly bill; keep the second model only if the measured gain clears your threshold.
For screenshots, PDFs, or document images, include Gemini in the candidate set because Google documents native multimodal support. Compare extraction accuracy, spatial grounding, Arabic-script handling, citations, latency, and current provider-listed cost on your own corpus before expanding the stack.
For deeper context: the comparison tool lets you pick any of these models and any dimension to compare, with a downloadable PDF. The cost guide covers pricing dynamics in detail.