Worth reading
DeepSeek-V4.1-Flash
DeepSeek · Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate
What people built this week, how to do it yourself, what to use, and what to do when it breaks. Every claim names its source and its date.
Someone did this
The review step was handled by the agents too, which is the part that usually does not survive contact with a deadline.
Reported by elder-pliniusSeptember 1, 2026Not tested by benchr
Which model, what it can do, what it costs, and what changed. Every figure names the page it was read from and the date.
One clear model decision
Three steps. No universal winner.
Worth reading
DeepSeek · Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate
New and changed
DeepSeek · Replacing the retired DeepSeek V4-Flash on the same API, at a lower listed rate
OpenAI · Long end-to-end work OpenAI aims this model at - reasoning, coding, computer use, research
Gemini 3.8 Flash is priced to the cent like 3.7 Flash. Google says it spends more tokens on purpose, which is where the cost difference lives.
Claude Fable 5.1 keeps the $10/$50 API rate and 1M context, cuts cache reads to $0.25 per 1M tokens, and changes forced-tool and thinking behavior.
Alibaba (Qwen) · High-volume multimodal work - among the lowest published rates benchr records
Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.
Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.
Grok 4.6 targets long coding agents, starts at $2/$6, doubles at 200K prompt tokens, and has no numeric text-output cap.
Anthropic's current Opus at the same $5/$25 rate — what changes over 4.8, and who should move.
A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.
A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.
Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.
ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.
Advertised context capacity versus reproducible retrieval checks you can run on your own documents.
The cheapest model for chat, coding, RAG, agents, classification, and summarization.
Where Llama 4, Mistral Large 3, DeepSeek-V4, and Qwen 3.6 stand against the closed labs.
MMLU is saturated above 90%. The benchmarks worth tracking now.
Claude Opus 5 and Fable 5.1, GPT-5.6 Sol and GPT-6 Astra, Gemini 3.1 Pro and Grok 4.6: sourced specifications, editorial tradeoffs, and a local evaluation plan.
Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.
What AI costs by model and workload, and where teams overspend.
An interactive comparison covering pricing, benchmarks, context windows, and capability ratings for all 38 models in the shared index, across frontier, mid, and open-weight tiers.
All 38 models ranked by benchr Rating →
All 114 articles in the archive →
Provider guides: OpenAI, Anthropic, Google, and open weights →
benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.