Corrections

Updates and corrections to published articles, with date and what changed.

This page lists material corrections made to articles after publication. Categories are factual (a number, name, or date was wrong), material (a verdict or recommendation changed), and stale (the underlying facts changed). Read the full editorial standards for how each category is handled.

If you spotted something we got wrong, send a note to corrections@benchr.org.

2026

September 12, 2026 — The open-weight coding leader was named wrong

Factual. The best free coding model guide ranked Kimi K2.6 first, calling it “the strongest by the numbers” at 80.2 on SWE-bench Verified. benchr’s own ledger already recorded DeepSeek-V4-Pro — MIT-licensed open weights, released April 24, 2026 — at 80.6 on the same benchmark, taken from DeepSeek’s model card. The ranking was wrong when it was published; it was not overtaken later. Both figures are provider-reported, and benchr has not run SWE-bench itself. The page now leads with DeepSeek-V4-Pro at 80.6, keeps Kimi K2.6 at 80.2, and states that the newer Kimi K3 publishes no SWE-bench Verified figure and so cannot be ranked on it. Affected surface: /articles/best-free-coding-model and its Arabic twin, from publication on May 30, 2026 until this correction.

September 5, 2026 — Claude Sonnet 5 carried a cancelled price increase

Factual. Anthropic announced that Claude Sonnet 5 would rise from $2/$10 to $3/$15 per 1M tokens on September 1, 2026, then cancelled that increase; benchr recorded the cancellation in the change ledger on August 31 but left the schedule in place in two data files. models.json and model-figures.json both kept standard_input_per_million_from_2026_09_01: 3.0, and the notes string that the model page prints verbatim still stated the rise as coming. The schedule and the notes are corrected; $2/$10 with cache-hit input at $0.20 is the standard price, read from Anthropic's pricing page on 2026-08-31. Affected surfaces: /pricing, /pricing/claude-sonnet-5, /models/claude-sonnet-5, and every consumer of the published models.json dataset, for the four days between the effective date and this correction.

September 5, 2026 — Gemini 3.8 Flash listed an introductory price as the price

Factual. benchr published $0.75/$3.75 per 1M tokens for Gemini 3.8 Flash with nothing to say the rate was introductory. Google's own announcement and pricing page, both read 2026-09-03, state that the introductory rate runs through December 31, 2026 and doubles to $1.50/$7.50 on January 1, 2027. The record now carries the end date and the standard rate, and the pricing table prints them. Anyone who budgeted the model from benchr between its September 2 release and this correction would have been wrong by a factor of two from January. Affected surfaces: /pricing and the published dataset.

July 30, 2026 — Qwen3.7-Max availability date

Factual. The recent-releases page and central model record incorrectly called May 19 the Qwen3.7-Max release date. Alibaba's own announcement is dated May 21 and said Model Studio API access was still “coming soon.” Current Model Studio documentation confirms that the API is available now, but does not state its first public availability date. We removed the unsupported date, recorded May 21 as the announcement date, and now distinguish current availability from an unknown first-availability date.

July 30, 2026 — Preview dates recorded as releases

Factual. The central record treated the announcement dates for Qwen3.6-Max-Preview and Muse Video Preview as product release dates. Alibaba's April 22 post contains conflicting “coming soon” and “available” language for Qwen3.6-Max-Preview, while Meta's July 7 post explicitly says Muse Video is coming soon. Both dates are now stored as announcement dates; no first-availability date is claimed without a provider source.

July 29, 2026 — Qwen AgentWorld model record

Factual. Removed the unsupported Qwen-AgentWorld-397B-A17B record. The official Qwen-AgentWorld repository documents the 35B-A3B open release, while the previously cited model URL did not provide a public record for the claimed 397B-A17B model.

August 21, 2026 — Publication provenance on 18 launch-day articles

Provenance. These pages were published on May 30, 2026, as their structured metadata states. Earlier changelog dates described retrospective subject snapshots or pre-publication preparation, but several entries incorrectly called those dates the original publication or an update. The English and Arabic changelogs now distinguish the May 30 publication date from the earlier coverage dates.

July 8, 2026 — Gemini 3.5 Pro review and pricing pages

Factual. The June 30, 2026 claim that Google released "Gemini 3.5 Pro" was fabricated. Both source URLs originally cited for the launch and its benchmark figures now return 404, Google's own Gemini API pricing and model documentation still list Gemini 3.1 Pro as the current flagship, and Google DeepMind's model hub lists "Gemini 3.5 Pro" as coming soon, not shipped. The review and pricing page have been withdrawn and replaced with correction notices; see the correction and the pricing correction.

July 8, 2026 — GPT-5.6 launch coverage

Factual. The July 1, 2026 claim that GPT-5.6 (Sol, Terra, Luna) had already reached general availability was false when checked July 8: OpenAI's developer docs still described a select-partner preview. OpenAI then confirmed the actual API release on July 9, 2026. The current launch article and pricing page preserve both dates so the correction is not mistaken for the current status.

July 8, 2026 — Claude Sonnet 5 max output

Factual. Claude Sonnet 5's max output was recorded as 200,000 tokens. The correct figure, per Anthropic's Models overview, is 128,000 tokens — which ties Claude Opus 4.8 and is double Sonnet 4.6's 64,000-token limit. Corrected across the review, the pricing page, and every comparison that cited the old figure.

May 25, 2026 — Multiple articles (pre-publication preparation)

Pre-publication record. Prepared Gemini lineup coverage to reflect the May 19, 2026 release of Gemini 3.5 Flash and the announcement of Gemini 3.5 Pro for June 2026. References were prepared to match the Google model lineup represented at publication.

May 4, 2026 — Pricing tables (retrospective research note)

Publication provenance. The May 30 launch version incorporated a pricing review against provider pages, including Claude Opus 4.7's April 16, 2026 launch rates.

January 22, 2026 — Open-weight tier (retrospective research note)

Publication provenance. The May 30 launch version included a retrospective open-weight market snapshot for this date; it was not a public January correction.

January 22, 2026 — Voice models compared (retrospective research note)

Publication provenance. The May 30 launch version included a retrospective voice-model snapshot for this date; it was not a public January correction.

January 15, 2026 — Coding assistants shootout (retrospective subject note)

Publication provenance. The May 30 launch version recorded the mid-January Cody and Windsurf updates as subject history; it was not a public January update.

January 8, 2026 — GPT-5, reviewed (retrospective subject note)

Publication provenance. The May 30 launch version discussed OpenAI's January API update as subject history; it was not a public January update.

Retrospective pre-launch notes

November 28, 2025 — GPT-5 vs Claude Opus 4.7 (retrospective research note)

Publication provenance. The May 30 launch version incorporated the corrected $10 GPT-5 output price for this retrospective comparison; it was not a public 2025 correction.