Model intelligence · sourced Source → shortlist → decision
Open workspace →

See what AI can really do — and how people are doing it

What people built this week, how to do it yourself, what to use, and what to do when it breaks. Every claim names its source and its date.

What people are making

All discoveries →

Things you can do with AI

Start from a goal →

Something not working?

All fixes →

What providers changed

All changes →

Explore the record

Open Explore →

Which model, what it can do, what it costs, and what changed. Every figure names the page it was read from and the date.

One clear model decision

Compare → test → choose.

Three steps. No universal winner.

38verified models
25logged changes
41retirement records
4 labelsevidence categories

Worth reading

DeepSeek-V4.1-Flash

DeepSeek · Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate

Model review

New and changed

Recent model updates

View all changes →
LoadingReading the verified change record…
  1. 01

    DeepSeek-V4.1-Flash

    DeepSeek · Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate

  2. 02

    GPT-6 Astra

    OpenAI · Long end-to-end work OpenAI aims this model at - reasoning, coding, computer use, research

  3. 03

    Gemini 3.8 Flash: same rate card, different bill

    Gemini 3.8 Flash is priced to the cent like 3.7 Flash. Google says it spends more tokens on purpose, which is where the cost difference lives.

  4. 04

    Claude Fable 5.1: cheaper cache reads, stricter contract

    Claude Fable 5.1 keeps the $10/$50 API rate and 1M context, cuts cache reads to $0.25 per 1M tokens, and changes forced-tool and thinking behavior.

  5. 05

    Qwen3.8-Flash

    Alibaba (Qwen) · High-volume multimodal work - among the lowest published rates benchr records

  6. 06

    GLM-5.3: 1M coding, hosted now with weights pending

    Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.

  7. 07

    Gemini 3.7 Flash: introductory pricing and migration

    Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.

  8. 08

    Grok 4.6: 500K agent model with no numeric output ceiling

    Grok 4.6 targets long coding agents, starts at $2/$6, doubles at 200K prompt tokens, and has no numeric text-output cap.

Reviews

  1. 07

    Claude Opus 5

    Anthropic's current Opus at the same $5/$25 rate — what changes over 4.8, and who should move.

  2. 08

    GPT-5

    A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.

  3. 09

    Gemini 3 Pro

    A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.

Comparisons

  1. 10

    Coding assistants

    Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.

  2. 11

    Voice models

    ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.

  3. 12

    Context windows

    Advertised context capacity versus reproducible retrieval checks you can run on your own documents.

Analysis

  1. 13

    Price per use case

    The cheapest model for chat, coding, RAG, agents, classification, and summarization.

  2. 14

    The open-weight tier

    Where Llama 4, Mistral Large 3, DeepSeek-V4, and Qwen 3.6 stand against the closed labs.

  3. 15

    Why benchmarks stopped telling you

    MMLU is saturated above 90%. The benchmarks worth tracking now.

Guides

  1. 16

    Frontier models

    Claude Opus 5 and Fable 5.1, GPT-5.6 Sol and GPT-6 Astra, Gemini 3.1 Pro and Grok 4.6: sourced specifications, editorial tradeoffs, and a local evaluation plan.

  2. 17

    Open-weight models

    Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.

  3. 18

    AI costs

    What AI costs by model and workload, and where teams overspend.

Compare models directly

An interactive comparison covering pricing, benchmarks, context windows, and capability ratings for all 38 models in the shared index, across frontier, mid, and open-weight tiers.

Open the comparison tool

All 38 models ranked by benchr Rating →

All 114 articles in the archive →

Recent model releases →

Provider guides: OpenAI, Anthropic, Google, and open weights →

API model ID directory and developer data →

benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.

Updates

Follow new pieces through RSS, recent releases, and the changelog.