Frontier AI models in 2026: a guide

A provider-sourced shortlist of hosted frontier models, plus a reproducible way to choose for your own workload.

By the benchr team ·

What this guide covers

This guide covers the hosted frontier candidates as of September 2026: Anthropic's Claude Opus 5 and Claude Fable 5.1, OpenAI's GPT-5.6 Sol and GPT-6 Astra, Google's Gemini 3.1 Pro (Google still labels it Preview), and xAI's Grok 4.6, with DeepSeek-V4-Pro and Kimi K3 as open-weight checks. The reviews and the comparison below were written for the earlier flagships — Claude Opus 4.7 and GPT-5 — and stay as the record of those models; their evaluation method carries over unchanged. The linked pages separate provider-published specifications and benchmarks from editorial hypotheses, and none claims an unpublished benchr test or a universal winner.

The current shortlist

Provider-listed API prices per 1M tokens (input / output) and context windows, as recorded by benchr on September 11, 2026. Benchmark figures are provider-reported.
ModelPrice / 1MContextWhat the provider publishes
Claude Opus 5$5 / $251MAnthropic's starting point for most workloads; benchr records no Anthropic benchmark table for it.
Claude Fable 5.1$10 / $501MAnthropic points demanding reasoning and long-horizon agentic work here. The API returns HTTP 400 when tool_choice is any or a named tool.
GPT-6 Astra$10 / $501.05MOpenAI's most capable model; no benchmark table published. Tool calling requires the Responses API.
GPT-5.6 Sol$4 / $201.05MSWE-bench Verified 89.8, GPQA Diamond 91.2 (OpenAI-reported).
Gemini 3.1 Pro$2 / $12 up to 200K-token prompts; $4 / $18 above~1MGPQA Diamond 94.3, SWE-bench Verified 80.6 (Google-reported). Still Google's newest Pro model.
Grok 4.6$2 / $6 under 200K-token prompts; $4 / $12 at or above500KxAI's flagship; benchr records no provider benchmark figure for it.
DeepSeek-V4-Pro (open weights, MIT)$1.32 / $3.96 peak; $0.66 / $1.98 off-peak1MSWE-bench Verified 80.6, Terminal-Bench 2.1 87.9 (DeepSeek-reported).
Kimi K3 (Kimi K3 License)$3 / $151MTerminal-Bench 2.1 88.3, GPQA 93.5 (Moonshot-reported); no SWE-bench Verified figure.

Reviews

  • Review · Nov 2025

    Claude Opus 4.7, reviewed

    Anthropic positions Opus 4.7 for demanding coding and agentic work and publishes its price and benchmark sheet. Treat architectural fit as an editorial hypothesis: test it on held-out repository issues and document tasks, then compare errors and review time.

  • Review · Jan 2026

    GPT-5, reviewed

    OpenAI publishes GPT-5 specifications and benchmark results that make it a candidate for structured output and reasoning workloads. Speed, prose preference, and technical error rate depend on the deployment and prompt; measure them on the same local rubric.

  • Review · Dec 2025

    Gemini 3.1 Pro, reviewed

    Google documents the model's context capacity and Workspace integrations. Image-heavy and Workspace-bound work are reasonable editorial hypotheses to test with representative screenshots, documents, citations, and failure cases.

Comparisons

  • Comparison · Dec 2025

    GPT-5 vs Claude Opus 4.7: a seven-task evaluation plan

    Seven workload categories with a shared prompt and scoring template. The category leans are editorial hypotheses, not private measured results; reproduce them with pinned versions, retained outputs, and blind review. Written for the earlier pair, the plan applies unchanged to Claude Opus 5 and GPT-5.6 Sol.

  • Comparison · Mar 2026

    Multimodal evaluation plan: twelve images, four models

    A reproducible rubric for Claude, GPT-5, Gemini 3, and Llama 4 across dense interfaces, document images, charts, and Arabic script. Placements are hypotheses to validate, not an unpublished winner table.

  • Analysis · Apr 2026

    The price-per-use-case table

    Six workload shapes using dated, provider-listed token prices. Recalculate with your input/output mix, cache hit rate, tool calls, retries, and current price pages before making a budget decision.

Which one should you use?

Start with the candidate whose provider-documented capabilities match your dominant workload. For coding and agentic work, put GPT-5.6 Sol (SWE-bench Verified 89.8, OpenAI-reported) next to Claude Opus 5, which Anthropic names as the model to start with for most workloads, though benchr records no Anthropic benchmark table for it. For multimodal, long-document, or Workspace-bound work, test Gemini 3.1 Pro (GPQA Diamond 94.3, Google-reported). Add Grok 4.6 when output cost dominates: at $2 / $6 per 1M tokens under 200K-token prompts, it has the lowest output rate among the closed models on this shortlist. These are editorial starting hypotheses, not measured winners.

Reserve GPT-6 Astra and Claude Fable 5.1, both $10 / $50 per 1M tokens, for tasks where your own evaluation shows the ceiling pays for itself. Check the integration first: GPT-6 Astra requires the Responses API for tool calling and accepts no custom temperature, and Claude Fable 5.1 returns HTTP 400 if tool_choice is any or a named tool. If the weights must stay on your own hardware, use DeepSeek-V4-Pro (MIT license, SWE-bench Verified 80.6, DeepSeek-reported) as the open-weight control.

To decide whether two subscriptions add value, route the same held-out set through each model alone and through the proposed pair. Record incremental task coverage, review time, failures, and the actual monthly bill; keep the second model only if the measured gain clears your threshold.

For screenshots, PDFs, or document images, include Gemini in the candidate set because Google documents native multimodal support. Compare extraction accuracy, spatial grounding, Arabic-script handling, citations, latency, and current provider-listed cost on your own corpus before expanding the stack.

For deeper context: the comparison tool lets you pick any of these models and any dimension to compare, with a downloadable PDF. The cost guide covers pricing dynamics in detail.