Antigravity Agent 05-2026
September 17, 2026Google announced the deprecation of Antigravity Agent 05-2026. Official replacement: Antigravity Agent 09-2026.
Google Tier 1
A projection over benchr's existing pricing history and lifecycle record, typed and linked to the capability graph. A change with no official source is not published here.
These records describe what providers officially document. benchr has not run these capabilities itself, so nothing here is a test result. Each record shows the vendor page it was read from and the date.
Ledger updated: September 3, 2026
Google announced the deprecation of Antigravity Agent 05-2026. Official replacement: Antigravity Agent 09-2026.
Google Tier 1
DeepSeek retired DeepSeek-V4-Flash and V4-Flash-Vision-Exp. Official replacement: DeepSeek-V4.1-Flash.
DeepSeek Tier 1
OpenAI announced gpt-6-astra in its API changelog on September 3, 2026 as its most capable model, for reasoning, coding, computer use, research and document creation. The model directory lists a 1.05M context window and an April 30, 2026 knowledge cutoff, and publishes no maximum-output figure. Pricing is $10.00 input / $1.00 cached input / $50.00 output per 1M tokens, with Batch at half of each. The changelog records interface constraints that are migration blockers: no `none` reasoning-effort level, no custom temperature, top_p or logprobs, and tool calling only through the Responses API. Source: developers.openai.com/api/docs/changelog, /pricing and /models (verified 2026-09-08).
OpenAI Tier 1
Google released gemini-3.8-flash on September 2, 2026, its third Flash release in six weeks. Official docs list a 1,048,576-token input limit, 65,536 max output, and tunable thinking levels. Introductory pricing is $0.75/$3.75 per 1M input/output through December 31, 2026, doubling to $1.50/$7.50 on January 1, 2027; batch and flex are $0.375/$1.875, priority $1.35/$6.75, cached input $0.075. Google published HLE-Verified 54.9% and a 47.2% pass@1 patching figure in the announcement, neither of which maps to a benchmark column benchr tracks, so no benchmark value is recorded. The 3.8 Flash Cyber variant is Fairwind Program only and is not separately priced. Google states 3.7 Flash remains fully supported. Sources: blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ and ai.google.dev/gemini-api/docs/pricing (verified 2026-09-03).
Google Tier 1
Anthropic released claude-fable-5-1 on September 1, 2026. Official docs list 1M context, 128K max output, $10/$50 standard input/output per 1M, $0.25 cached input, and Batch at $5/$25. Adaptive thinking is always on. Source: anthropic.com/claude-fable-and-mythos-5-1 and platform.claude.com/docs/en/models/fable-5-1/overview (verified 2026-09-02).
Anthropic Tier 1
Moonshot AI retired Kimi K2.5 and Moonshot V1 series. Official replacement: Kimi K3.
Moonshot AI Tier 1
Anthropic's pricing page now lists $2/$10 per 1M as Claude Sonnet 5's standard price. The increase to $3/$15 scheduled for September 1, 2026 was cancelled; cache-hit input stays $0.20. Verified 2026-08-31.
Provider changes affecting this: One screenshot in, a working page out · Drive a real browser, not a scraper · Operate a desktop application · Understand a codebase nobody has explained to you · Analyze a data file and get a chart you can trust
Anthropic Tier 1
Cache-miss input rose from a flat $0.435 to $1.32 peak / $0.66 off-peak per 1M tokens, output from $0.87 to $3.96 / $1.98, and cache-hit input from $0.003625 to $0.044 / $0.022. DeepSeek replaced flat API pricing with peak/off-peak billing at 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; all other hours bill at half the peak rate. The pricing figures here are the peak (list) rates. Source: api-docs.deepseek.com/quick_start/pricing (read 2026-08-28).
DeepSeek Tier 1
Cache-miss input rose from a flat $0.14 to $0.44 peak / $0.22 off-peak per 1M tokens, output from $0.28 to $1.32 / $0.66, and cache-hit input from $0.0028 to $0.014 / $0.007. DeepSeek replaced flat API pricing with peak/off-peak billing at 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; all other hours bill at half the peak rate. The pricing figures here are the peak (list) rates. Source: api-docs.deepseek.com/quick_start/pricing (read 2026-08-28).
DeepSeek Tier 1
Alibaba released Qwen3.8-Flash on August 27, 2026: a multimodal mixture-of-experts model with a 125B main model plus 51B of N-gram embeddings, activating 6B parameters per token, with a native 262K context the announcement says extends to 1M. Model Studio lists $0.15 input / $0.47 output per 1M tokens in the Singapore region. Alibaba names SWE-bench Pro, CoWorkBench, Toolathlon Verified, MathVision, AndroidWorld and ERQA but publishes no scores, and names no license. Source: alibabacloud.com/blog/...603503 and alibabacloud.com/help/en/model-studio/model-pricing (verified 2026-09-08).
Alibaba (Qwen) Tier 1
OpenAI retired OpenAI Assistants API. Official replacement: Responses API + Conversations API.
OpenAI Tier 1
Google's pricing table now lists Gemini 3.6 Flash at $0.75/$3.75 per 1M tokens through December 31, 2026, down from the $1.50/$7.50 recorded on 2026-07-22, with $1.50/$7.50 returning January 1, 2027. Context caching drops to $0.075. Source: ai.google.dev/gemini-api/docs/pricing (read 2026-08-24).
Google Tier 1
Moonshot released kimi-k3 on 2026-07-16 and later published official weights. Global hosted pricing is $3 cache-miss input, $0.30 cache-hit input, and $15 output per 1M; the model card lists 2.8T total and 104B active parameters. Source: kimi.ai technical blog, platform.kimi.ai docs, and the official Moonshot AI model card (verified 2026-08-21).
Moonshot AI Tier 1
xAI launched grok-4.6 on 2026-08-12. Official pricing lists $2 input / $0.50 cached / $6 output per 1M below 200K prompt tokens, and $4 / $1 / $12 at or above 200K; the higher tier applies to all tokens in that request. The developer guide lists 500K context and no numeric text-output limit. Sources: https://x.ai/news/grok-4-6, https://docs.x.ai/developers/pricing, and https://docs.x.ai/developers/grok-4-6 (verified 2026-08-21).
xAI Tier 1
Z.AI released the hosted glm-5.3 endpoint on 2026-08-18 with $1.40/$4.40 per 1M, $0.26 cached input, 1M context, and 128K max output. Weights were announced for later release after safety hardening and were not recorded as available. Source: z.ai/blog/glm-5.3 and docs.z.ai (verified 2026-08-21).
Z.ai Tier 1
Google released gemini-3.7-flash as stable GA on 2026-08-13. Introductory Standard pricing through 2026-12-31 is $0.75/$3.75 per 1M with $0.075 cached input; scheduled 2027 Standard pricing is $1.50/$7.50. Source: ai.google.dev Gemini API changelog, model, latest-model, and pricing pages (verified 2026-08-21).
Provider changes affecting this: Let the model write and run real code mid-answer · Fit a whole codebase in one prompt, and find its limit · Understand a codebase nobody has explained to you · Analyze a data file and get a chart you can trust
Google Tier 1
Anthropic retired Anthropic experimental prompt tools API and Workbench.
Anthropic Tier 1
OpenAI retired GPT-5.2 / 5.3 chat-latest snapshots. Official replacement: GPT-5.6 Sol.
OpenAI Tier 1
Google retired Gemini Embedding 2 Preview. Official replacement: Gemini Embedding 2.
Google Tier 1
Anthropic retired Claude Opus 4.1. Official replacement: Claude Opus 4.8.
Anthropic Tier 1
Alibaba announced Qwen3.8-Max on August 3, 2026 as its largest flagship: a sparse mixture-of-experts model with 2.4T total parameters activating 95B per token, multimodal across text and vision, with a context window described as up to 1M tokens. Model Studio lists $2.00 input / $6.00 output per 1M tokens in the Singapore region. Alibaba published arena placements rather than numeric scores - fifth in Text Arena, second in Vision Arena, fourth in Frontend Code Arena - and a placement is not a score. Weights were said to follow the next week; no license was verified on an official page. Source: alibabacloud.com/blog/...603420 and alibabacloud.com/help/en/model-studio/model-pricing (verified 2026-09-08).
Alibaba (Qwen) Tier 1
Google announced the deprecation of Gemini Robotics ER 1.6 Preview. Official replacement: Gemini Robotics ER 2 Preview.
Google Tier 1
DeepSeek retired deepseek-chat / deepseek-reasoner (legacy aliases). Official replacement: DeepSeek V4-Flash / DeepSeek V4-Pro.
DeepSeek Tier 1
Anthropic retired Claude Opus 4.7 fast mode. Official replacement: Claude Opus 4.8 fast mode or Claude Opus 4.7 standard mode.
Anthropic Tier 1
OpenAI retired GPT-5 era Codex models. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI retired o4-mini-deep-research. Official replacement: GPT-5.5 Pro.
OpenAI Tier 1
OpenAI retired o3-deep-research. Official replacement: GPT-5.5 Pro.
OpenAI Tier 1
OpenAI retired GPT-5 / 5.1 chat-latest snapshots. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI retired computer-use-preview. Official replacement: GPT-5.4 mini.
OpenAI Tier 1
OpenAI announced the deprecation of OpenAI legacy audio, realtime, and transcription models. Official replacement: GPT Realtime 2.1 / GPT Audio 1.5.
OpenAI Tier 1
Correction after live re-read of Anthropic's official pricing page: Claude Sonnet 5 launches with introductory API pricing of $2 input / $10 output per 1M tokens through August 31, 2026, then $3/$15 from September 1, 2026. Previous benchr July 1 entry incorrectly recorded $4/$20. Source: platform.claude.com/docs/en/about-claude/pricing
Provider changes affecting this: One screenshot in, a working page out · Drive a real browser, not a scraper · Operate a desktop application · Understand a codebase nobody has explained to you · Analyze a data file and get a chart you can trust
Anthropic Tier 1
Released July 1, 2026. Anthropic's second Mythos-class-architecture model after Fable 5, priced mid-tier between Sonnet 4.6 and Opus 4.8; 128,000-token max output. Safety-classified requests return an explicit refusal; another-model retry requires configured application logic. Shipped the same day Anthropic restored Claude Fable 5 to all customers. Source: anthropic.com/news/claude-sonnet-5
Provider changes affecting this: One screenshot in, a working page out · Drive a real browser, not a scraper · Operate a desktop application · Understand a codebase nobody has explained to you · Analyze a data file and get a chart you can trust
Anthropic Tier 1
Google retired Gemini 3 Pro Image Preview. Official replacement: Gemini 3 Pro Image (stable).
Google Tier 1
Google retired Gemini 3.1 Flash Image Preview. Official replacement: Gemini 3.1 Flash Image (stable).
Google Tier 1
Google announced the deprecation of Imagen 4.0 image models. Official replacement: Gemini 3.1 Flash Image.
Google Tier 1
Anthropic retired Claude Sonnet 4. Official replacement: Claude Sonnet 4.6.
Anthropic Tier 1
Anthropic retired Claude Opus 4. Official replacement: Claude Opus 4.8.
Anthropic Tier 1
OpenAI announced the deprecation of Older GPT-5 and o3 model snapshots. Official replacement: GPT-5.5 / GPT-5.4 mini / GPT-5.4 nano / GPT-5.5 Pro.
OpenAI Tier 1
Re-added after verification. Released March 5, 2026; wrongly removed June 1 as 'unverified'. Source: openai.com/api/pricing
OpenAI Tier 1
Released June 9, 2026. Mythos-class and generally available; safety-classified cyber/bio/distillation requests return an explicit refusal, and another-model retry requires configured application logic. Suspended June 12 and restored globally July 1, 2026. Source: anthropic.com/news/claude-fable-5-mythos-5
Provider changes affecting this: Drive a real browser, not a scraper · Operate a desktop application · Stop paying twice for the same file · Automate something on a site that has no API · Cut the API bill without changing the answer
Anthropic Tier 1
OpenAI announced the deprecation of OpenAI Reusable Prompts API.
OpenAI Tier 1
OpenAI announced the deprecation of OpenAI Evals platform (dashboard + API).
OpenAI Tier 1
OpenAI announced the deprecation of OpenAI Agent Builder. Official replacement: Agents SDK.
OpenAI Tier 1
OpenAI announced the deprecation of GPT Image 1 mini and 1.5. Official replacement: GPT Image 2.
OpenAI Tier 1
Mistral Large 3 released December 2, 2025. Apache 2.0. API id: mistral-large-2512. Source: mistral.ai/technology/models
Mistral AI Tier 1
Llama 4 Scout released April 5, 2025. Community License. 10M token context. Self-hosted. Source: ai.meta.com/llama
Provider changes affecting this: Run a capable model on your own machine · Cut the API bill without changing the answer · Use a model without sending anything anywhere
Meta Tier 1
Llama 4 Maverick released April 5, 2025. Community License. Self-hosted, no API fee. Source: ai.meta.com/llama
Meta Tier 1
Kimi K2.6 released April 20, 2026. Modified MIT. Source: platform.moonshot.ai/docs/pricing
Moonshot AI Tier 1
Grok 4.3 — $1.25/$2.50; cached $0.20. Source: x.ai/api
xAI Tier 1
GPT-5 launched August 7, 2025. $1.25/$10. Source: openai.com/pricing
Provider changes affecting this: Get JSON that always matches your schema · Analyze a data file and get a chart you can trust · Check whether the answer is actually right
OpenAI Tier 1
GPT-5 Mini — $0.25/$2.00. Cached input $0.025. Source: openai.com/pricing
OpenAI Tier 1
GPT-5.5 launched April 24, 2026. $5/$30 standard; $2.50/$15 batch. Source: openai.com/pricing
Provider changes affecting this: Get JSON that always matches your schema · Analyze a data file and get a chart you can trust · Check whether the answer is actually right
OpenAI Tier 1
Gemini 3.5 Flash GA since May 19, 2026 (Google I/O). $1.50/$9.00. Source: ai.google.dev/pricing
Provider changes affecting this: Let the model write and run real code mid-answer · Analyze a data file and get a chart you can trust · Build a working app from an idea
Google Tier 1
Google retired Gemini 2.0 Flash family. Official replacement: Gemini 3.5 Flash.
Google Tier 1
DeepSeek V4-Pro launched April 24, 2026. MIT license. Source: api-docs.deepseek.com/quick_start/pricing
DeepSeek Tier 1
DeepSeek V4-Flash launched April 24, 2026. MIT license. Cheapest commercial API in this set. Source: api-docs.deepseek.com
DeepSeek Tier 1
Claude Opus 4.8 launched May 28, 2026. $5/$25 standard; $10/$50 fast mode. SWE-Bench Pro 69.2%. 1M context window confirmed on platform.claude.com. Source: anthropic.com/news
Provider changes affecting this: Drive a real browser, not a scraper · Operate a desktop application · Let the model write and run real code mid-answer · Understand a codebase nobody has explained to you · Analyze a data file and get a chart you can trust
Anthropic Tier 1
OpenAI retired DALL-E 2 and DALL-E 3. Official replacement: GPT Image 2.
OpenAI Tier 1
OpenAI announced the deprecation of o4-mini. Official replacement: GPT-5.4 mini.
OpenAI Tier 1
OpenAI announced the deprecation of o3-mini. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI announced the deprecation of o1-pro. Official replacement: GPT-5.5 Pro.
OpenAI Tier 1
OpenAI announced the deprecation of o1. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI announced the deprecation of GPT Image 1. Official replacement: GPT Image 2.
OpenAI Tier 1
OpenAI announced the deprecation of GPT-4o. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI announced the deprecation of GPT-4 Turbo. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI announced the deprecation of GPT-4.1 nano. Official replacement: GPT-5.4 nano.
OpenAI Tier 1
OpenAI announced the deprecation of GPT-4. Official replacement: GPT-5.5.
OpenAI Tier 1
OpenAI announced the deprecation of GPT-3.5 Turbo. Official replacement: GPT-5.4 mini.
OpenAI Tier 1
Anthropic retired Claude Haiku 3. Official replacement: Claude Haiku 4.5.
Anthropic Tier 1
OpenAI announced the deprecation of Sora 2 and Sora 2 Pro.
OpenAI Tier 1
OpenAI announced the deprecation of OpenAI Videos API.
OpenAI Tier 1
Google retired Gemini 3 Pro Preview. Official replacement: Gemini 3.1 Pro Preview.
Google Tier 1
Anthropic retired Claude Sonnet 3.7. Official replacement: Claude Sonnet 4.6.
Anthropic Tier 1
Anthropic retired Claude Haiku 3.5. Official replacement: Claude Haiku 4.5.
Anthropic Tier 1
Anthropic retired Claude 3 Opus. Official replacement: Claude Opus 4.8.
Anthropic Tier 1
Anthropic retired Claude Sonnet 3.5 (both snapshots). Official replacement: Claude Sonnet 4.6.
Anthropic Tier 1
Nothing matches that yet
The ledger holds 20 documented capabilities, so a narrow filter empties quickly. That is the honest state, not a search failure.
benchr records what providers changed. Whether it still does what your own work needs is a different question, and it is the one Labs answers.
Open Labs