# benchr > An AI model intelligence and decision platform: sourced model records, evaluation workspaces, production-cost scenarios, migration guidance, and editorial analysis for frontier and open-weight models. benchr is an editorial publication that synthesizes public information (official provider pricing pages, benchmark leaderboards, and model documentation) into reviews, head-to-head comparisons, and buying guidance for frontier and open-weight AI models. Provider facts link to primary sources and dates. Provider-published figures, third-party results, and benchr editorial estimates are identified separately; missing official figures may remain blank. Material corrections are public at /corrections, and commercial relationships are disclosed at /affiliate-disclosure. ## What AI can do, what people did with it, and what breaks - [Discover](https://benchr.org/discover): 12 things people actually did with AI, each attributed to whoever did it and linked to the original post. benchr has reproduced none of them; every record says so on its face. - [Do](https://benchr.org/do): 23 goal-first entry points. Each names the documented routes to the goal, ordered by effort, with the limits attached. - [Fix](https://benchr.org/fix): 16 things that go wrong in practice — what is happening, why, a quick fix, a real fix, and what no fix solves. Each record says whether the cause is documented by a vendor or is a reported pattern nobody documents. - [Techniques](https://benchr.org/techniques): 12 repeatable moves, each written as input, technique, output, with the failure modes named. benchr has run none of them under a published protocol. - [Learn](https://benchr.org/learn): 6 reading orders over material benchr already publishes. No new prose; every step points at an existing page. - [Explore](https://benchr.org/explore): The way into the reference layer — models, capabilities, pricing, comparisons, changes and retirements. - [Verified capabilities](https://benchr.org/capabilities): 20 concrete things AI can be made to do, each naming the vendor page it was read from and the date. - [Playbooks](https://benchr.org/workflows): 9 complete processes with a page each. Every stage names what you start with, what to do and what you end with, and says whether it is an AI step or human work. - [AI tools and surfaces](https://benchr.org/ai-tools): 7 surfaces a capability actually runs on; six have their own page. - [Verification runs](https://benchr.org/verification): 7 executed runs, published in full whether they passed or not. All are self-reported by the agent that maintains the repository, none is accepted, and therefore **no capability carries Benchr Verified**. **What "Verified" means here.** Two independent axes, never merged. `officially_supported` means benchr read the vendor's own current documentation on the date shown, and it said what the record says. `benchr_verified` means benchr ran the capability and an independent human reproduced it — and today **no record carries it**. Every capability reads `unverified` on the testing axis, every record carries an explicit `notVerified` list of what its source did not state, and nothing on /discover, /fix or /techniques is a benchr test result. ## Product platform - [Model intelligence platform](https://benchr.org/platform): Explore 34 models, inspect factual versus editorial evidence, shortlist candidates, model a workload and budget, save guest workspaces locally, watch lifecycle changes, and share URL state. - [Evaluation Labs](https://benchr.org/labs): Build or import test cases, compare 2–5 model outputs, score responses blind, add measured latency and token counts, and export or explicitly share a client-side evaluation. - [Prompt Workbench](https://benchr.org/prompt-workbench): Build an explicit prompt contract, capture version A, flag multi-section changes, export the draft, and hand one or two private cases to Labs through session storage rather than the URL. - [مساحة بناء البرومبت](https://benchr.org/ar/prompt-workbench): نسخة عربية أصلية لبناء عقد البرومبت، وضبط اختبار A/B، وتمرير الحالات محليًا إلى المختبر. - [Use-case evaluation packs](https://benchr.org/assets/data/use-case-eval-suites.json): Versioned original prompts and scoring rubrics used by Labs; the dataset contains no model outputs or benchmark results and does not claim that benchr ran the models. - [Migration assistant](https://benchr.org/migrate): Match exact retiring API IDs against the sourced lifecycle record, separate official replacements from benchr alternatives, compare recorded prices, and generate a production migration checklist. - [Production cost calculator](https://benchr.org/calculator): Compare monthly token cost across 34 models with prompt caching, batch pricing, and reasoning-output scenarios; export CSV or share the scenario URL. - [Arabic platform](https://benchr.org/ar/platform): RTL Arabic edition of the model explorer and local decision workspace. - [Local AI reference](https://benchr.org/local): Source-linked local open-weight model records, transparent Q4 weight planning, and an interactive “What can I run?” tool. The result is a planning label, not a performance guarantee. - [دليل الذكاء الاصطناعي المحلي بالعربية](https://benchr.org/ar/local): نسخة عربية أصلية من مرجع النماذج المحلية ومخطط الذاكرة، مع نفس حدود التخطيط وروابط المصادر. - [What can I run? memory planner](https://benchr.org/local/what-can-i-run): Interactive local-AI memory planning with published formulas and explicit headroom policy. ## Enterprise agents - [OpenAI Presence: what the enterprise agent service actually includes](https://benchr.org/articles/openai-presence-enterprise-agents): A procurement-focused analysis of the managed voice-and-chat product, its control and evaluation layers, provider-reported outcomes, missing public price, and a pilot checklist. - [OpenAI Presence — Arabic edition](https://benchr.org/ar/articles/openai-presence-enterprise-agents): Independent Arabic buying guide with the same sourced facts and contract questions. ## Read-only API v1 - [API discovery](https://benchr.org/api/v1): Machine-readable endpoint index. - [Models API](https://benchr.org/api/v1/models): Current model records with explicit factual, provider, benchr, and editorial provenance. - [Deprecations API](https://benchr.org/api/v1/deprecations): Lifecycle status, API IDs, shutdown dates, sources, and replacement fields. - [Changes API](https://benchr.org/api/v1/changes): Append-only model history plus dated lifecycle announcements. - [Recommendations API](https://benchr.org/api/v1/recommendations): Transparent editorial heuristic with queryable task, budget, priority, privacy, and provider assumptions. - `POST /api/v1/inference` is optional and disabled unless the deployment owner explicitly enables it and configures a privacy acknowledgement, server-side OpenRouter key, model allowlist, and private bearer token; no secret is shipped to the browser. ## Guides - [Frontier AI models in 2026](https://benchr.org/guides/frontier-models): Claude Opus 5 and Fable 5.1, GPT-5.6 Sol and GPT-6 Astra, Gemini 3.1 Pro and Grok 4.6 — provider-published facts and a reproducible workload test plan. - [Open-weight AI models](https://benchr.org/guides/open-weight-models): Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host. - [AI costs in 2026](https://benchr.org/guides/ai-costs): What AI actually costs by model and workload, and where teams overspend. ## Reviews — Claude - [Claude Fable 5.1 migration guide](https://benchr.org/articles/claude-fable-5-1-migration): Anthropic's September 1 endpoint at 1M context and 128K output; recorded pricing, access boundary, and request-contract changes to test before upgrading. - [Claude Opus 4.8, reviewed](https://benchr.org/articles/claude-opus-4-8-review): A six-week upgrade at the same $5/$25 price, a 69.2% SWE-bench Pro score, and honesty gains that catch flaws in the model's own code. - [Claude Opus 4.7, reviewed](https://benchr.org/articles/claude-opus-4-7-review): Coding, long-document analysis, and multilingual capability at $5/$25 per million input/output tokens. - [Claude Sonnet 4.6, reviewed](https://benchr.org/articles/claude-sonnet-4-6-review): The $3/$15 daily-driver tier with a 1M-token window, when it's the right default, and when to pay for Opus. - [Claude Haiku 4.5, reviewed](https://benchr.org/articles/claude-haiku-4-5-review): The $1/$5 cost-control tier — where the cheapest current Claude is enough and how to use it in a routing strategy. - [Claude Mythos: the model you can't use](https://benchr.org/articles/claude-mythos): Anthropic's restricted frontier model, gated under Project Glasswing for cybersecurity — what it is and why it's locked away. - [Claude Fable 5 launch](https://benchr.org/articles/claude-fable-5-launch): A dated account of the Mythos-class launch, $10/$50 list pricing, safety classifiers, subsequent access changes, and the historical June free window. - [Claude Cowork: the desktop agent for non-coders](https://benchr.org/articles/claude-cowork): Claude Code's agentic engine for desktop knowledge work across local files and apps, plus the rollout. ## Reviews — Other providers - [GPT-5, reviewed](https://benchr.org/articles/gpt-5-review): Visual design, structured output, and language breadth, plus confident-but-wrong failure modes on niche technical questions. - [GPT-5.5, reviewed](https://benchr.org/articles/gpt-5-5-review): What it changes over GPT-5 on agentic coding and computer use, at roughly double the API price. Who should upgrade and who should wait. - [GPT-5.4, reviewed](https://benchr.org/articles/gpt-5-4-review): The seven-week flagship retrospective — $2.50/$15, 1M context, 75% OSWorld computer use, and why losing the crown made it the value pick. - [Gemini 3 Pro, reviewed](https://benchr.org/articles/gemini-3-pro-evaluation): Best-in-class vision and a 1M-token context window; average elsewhere. Deprecated March 2026 for Gemini 3.1 Pro Preview. - [Gemini 3.1 Pro, reviewed](https://benchr.org/articles/gemini-3-1-pro-review): A verified ARC-AGI-2 jump to 77.1%, a top GPQA Diamond score, 1M context, and a tiered long-context pricing cliff to watch. - [Gemini 3.7 Flash costs the same as 3.6 Flash](https://benchr.org/articles/gemini-3-7-flash-price-parity): Google's August 13, 2026 Flash release lists $0.75 input and $3.75 output per 1M tokens through December 31, 2026, the identical rate its predecessor was cut to, with both returning to $1.50/$7.50 on January 1, 2027 and no published benchmark table. - [Gemini 3.6 Flash launch](https://benchr.org/articles/gemini-3-6-flash-launch): Google's July 2026 stable Flash update, now priced at $0.75 input and $3.75 output through December 31, 2026, with migration guidance for older Flash endpoints. - [Gemini 3.5 Flash, reviewed](https://benchr.org/articles/gemini-3-5-flash-review): $1.50/$9 per 1M tokens, published capabilities, and clearly separated provider claims and editorial planning estimates. - [Grok 4.3, reviewed](https://benchr.org/articles/grok-4-3-review): Native live X and web search built into the API, $1.25/$2.50 pricing, a 1M-token window, and where its real-time edge does and doesn't win. - [Grok 4.5, reviewed](https://benchr.org/articles/grok-4-5-review): xAI's coding and agent specialist, built with Cursor, priced at $2/$6, and running a smaller 500K context window than Grok 4.3. - [GLM-5.2, reviewed](https://benchr.org/articles/glm-5-2-review): Zhipu's MIT-licensed, open-weight 753B model at $1.40/$4.40 per million tokens, with provider-reported coding benchmarks ahead of GPT-5.5 and Opus 4.7. - [DeepSeek-V4, reviewed](https://benchr.org/articles/deepseek-review): An MIT open-weight coding model with a 1M-token context and a hosted API that massively undercuts closed frontier output. - [DeepSeek peak and off-peak pricing](https://benchr.org/articles/deepseek-peak-pricing): DeepSeek replaced flat V4 API pricing on August 16, 2026 with peak rates for 01:00-04:00 and 06:00-10:00 UTC Monday to Friday and half those rates otherwise; V4-Flash went from $0.14/$0.28 to $0.44/$1.32 peak and $0.22/$0.66 off-peak. - [DeepSeek legacy API aliases retired](https://benchr.org/articles/deepseek-api-aliases-retired): A sourced July 24, 2026 migration checklist for replacing deepseek-chat and deepseek-reasoner with explicit V4 model IDs. - [Qwen3.6, reviewed](https://benchr.org/articles/qwen-review): An Apache-2.0 open-weight family in two sizes with up to a 1M context, free to self-host, built for agentic coding. - [Kimi K2.6, reviewed](https://benchr.org/articles/kimi-review): An open-weight trillion-parameter MoE for agentic and coding work, with an Agent Swarm of up to 300 sub-agents and a 256K context. - [Llama 4, reviewed](https://benchr.org/articles/llama-4-review): Scout's 10M-token context still stands out, but with Meta's flagship now closed-source, Llama 4 is the last major open-weight Llama. - [Mistral Large 3, reviewed](https://benchr.org/articles/mistral-review): An Apache-2.0 open-weight MoE flagship with a 256K context and multimodal, multilingual reach, plus Medium 3.5 and Small 4. - [ChatGPT Images 2.0, reviewed](https://benchr.org/articles/chatgpt-images-review): The image model that finally renders readable text in pictures — where it nails dense layouts and where it slips. ## Current model release records - [Claude Fable 5.1: cheaper cache reads, stricter contract](https://benchr.org/articles/claude-fable-5-1-migration): Anthropic's September 1 release adds a 1M-context, 128K-output endpoint at $10/$50 per million, with $0.25 cache reads and documented breaking request changes. - [Gemini 3.7 Flash: price the introductory window before you migrate](https://benchr.org/articles/gemini-3-7-flash-review): Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change. - [Grok 4.6: a 500K agent model whose output ceiling is not a number](https://benchr.org/articles/grok-4-6-review): xAI targets long coding runs and interactive agents at $2/$6, while its docs publish no numeric text-output cap. - [GLM-5.3: a 1M coding model with weights still behind a safety gate](https://benchr.org/articles/glm-5-3-review): Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40. - [Kimi K3: open weights at 2.8T parameters, but only 104B active](https://benchr.org/articles/kimi-k3-review): Moonshot's sparse agent model pairs a 1M context with flat hosted pricing and an open-weight deployment decision. - [Claude Opus 5: the migration question behind Anthropic's new flagship](https://benchr.org/articles/claude-opus-5-review): The same $5/$25 list price, a 1M-token window, and one configuration change that can break a careless upgrade. - [GPT-Realtime-2.1: evaluate the conversation loop, not a text-only price card](https://benchr.org/articles/gpt-realtime-2-1-review): OpenAI's speech-to-speech model adds reasoning and tool use, but its audio bill needs its own test plan. - [GPT-Realtime-2.1 mini: the lower-cost voice model still has an audio bill](https://benchr.org/articles/gpt-realtime-2-1-mini-review): OpenAI's distilled realtime reasoning tier keeps the large context window while changing the cost shape. - [GPT-Live-1 is a ChatGPT Voice rollout, not an API endpoint](https://benchr.org/articles/gpt-live-1-launch): OpenAI's full-duplex Voice model changes the product experience while leaving a developer pricing checklist empty. - [GPT-Live-1 mini: the Voice fallback is not a developer SKU](https://benchr.org/articles/gpt-live-1-mini-launch): OpenAI's smaller full-duplex ChatGPT Voice model reaches a different user tier, with no public API contract yet. - [Gemini 3.5 Flash-Lite: the $0.30 subagent target has a migration cost](https://benchr.org/articles/gemini-3-5-flash-lite-review): Google's lowest-cost 3.5 model is built for throughput, but the API changes are part of the selection decision. - [Gemini 3.1 Flash TTS Preview: steerable speech with preview boundaries](https://benchr.org/articles/gemini-3-1-flash-tts-preview-review): Google's text-to-speech model has explicit token limits and Batch support, but no public per-token price in the checked model card. - [Gemma 4 26B A4B IT: an official availability record, not a filled-in spec sheet](https://benchr.org/articles/gemma-4-26b-a4b-it-launch): Google published the model ID and route to use it. It did not publish the context, price, or benchmark figures this page would need to claim. - [Gemma 4 31B IT: a public model name without a public deployment spec](https://benchr.org/articles/gemma-4-31b-it-launch): Google confirmed the 31B IT identifier and availability. The missing limits, prices, and benchmark table must stay missing. - [Leanstral 1.5 is for proofs, not chat: read its benchmark claims in context](https://benchr.org/articles/leanstral-1-5-review): Mistral's open-weight formal-verification specialist has an unusual architecture and provider-reported proof metrics—not a general chatbot score. - [Qwen-Audio 3.0 TTS Plus: built-in voices are the product boundary](https://benchr.org/articles/qwen-audio-3-0-tts-plus-review): Alibaba's instruction-controlled speech model targets standard synthesis, with published availability but no public token limits or per-token price in the checked docs. - [Qwen-Audio 3.0 TTS Flash: latency and cloning change the deployment decision](https://benchr.org/articles/qwen-audio-3-0-tts-flash-review): Alibaba's low-latency TTS branch adds voice cloning, so consent and call-flow tests matter alongside speed. - [Muse Image: Meta's consumer launch leaves a developer checklist blank](https://benchr.org/articles/muse-image-launch): The image model is available in Meta AI, but the checked announcement gives developers no public model ID, rate card, or token limits. - [Muse Video Preview: native audio is a preview feature, not an API contract](https://benchr.org/articles/muse-video-preview-launch): Meta previewed a video model with native audio alongside Muse Image, while leaving public developer terms unannounced. ## Comparisons - [GPT-5 vs Claude Opus 4.7](https://benchr.org/articles/gpt-5-vs-claude-opus): Seven editorial decision factors grounded in cited public evidence, with no unpublished head-to-head score or universal winner. - [Opus 4.8 vs GPT-5.5: coder's pick vs daily driver](https://benchr.org/articles/opus-4-8-vs-gpt-5-5): Where Opus wins on coding, where it loses on Terminal-Bench, and which fits your stack at the same $5 input price. - [Gemini 3.1 Pro vs GPT-5.5](https://benchr.org/articles/gemini-3-1-pro-vs-gpt-5-5): Gemini chases hardest-mode reasoning at $2/$12; GPT-5.5 chases all-round knowledge work at $5/$30. - [Grok 4.3 vs ChatGPT: when live context wins](https://benchr.org/articles/grok-4-3-vs-chatgpt): Grok plugs into the live web and X; ChatGPT is the all-round assistant — a scenario-by-scenario guide. - [ChatGPT vs Claude vs Gemini: the 2026 pick](https://benchr.org/articles/chatgpt-vs-claude-vs-gemini): Three near-identical $20 subscriptions, scored on four everyday tasks, with a clear pick for each user. - [Claude vs ChatGPT for long-form writing](https://benchr.org/articles/claude-vs-chatgpt-writing): Output ceilings, voice, and instruction-following for long-form writing: how much each produces and which to trust. - [AI search engines: Perplexity vs ChatGPT vs Google](https://benchr.org/articles/ai-search-engines-compared): Perplexity, ChatGPT Search, and Google AI Overviews on sourcing, accuracy, and when to use each. - [The coding assistants shootout](https://benchr.org/articles/coding-assistants-shootout): Cursor, GitHub Copilot, Windsurf, and Cody on architecture, model backend, and bug profile. - [Multimodal capability ranking](https://benchr.org/articles/multimodal-capability-ranking): A source-based comparison of published vision capabilities across Claude, GPT-5, Gemini, and Llama, with editorial judgments labelled. - [Voice models compared](https://benchr.org/articles/voice-models-compared): ElevenLabs, OpenAI, and Cartesia compared using published pricing, latency claims, language support, and clearly labelled editorial trade-offs. - [Context windows compared](https://benchr.org/articles/context-windows-compared): The advertised window versus the effective retrieval zone where models reliably find information. - [AI model pricing comparison 2026](https://benchr.org/articles/ai-model-pricing-comparison): Cost per million tokens across OpenAI, Anthropic, Google, DeepSeek, and open-weight models. - [Cheapest LLM API 2026](https://benchr.org/articles/cheapest-llm-api-2026): Ultra-low-cost API models ranked by price, plus the trade-offs that come with the cheapest tier. - [Anthropic Claude API pricing guide](https://benchr.org/articles/claude-api-pricing-guide): Opus 4.8, Sonnet 4.6, and Haiku 4.5 pricing, including caching and batch discounts. - [DeepSeek vs OpenAI pricing](https://benchr.org/articles/deepseek-vs-openai-pricing): Cost comparison and quality trade-offs between DeepSeek and OpenAI pricing tiers. - [OpenAI API pricing guide](https://benchr.org/articles/openai-api-pricing-guide): GPT-5.5, GPT-5, and GPT-5 Mini pricing, including input, output, caching, and batch costs. ## Roundups — Best AI for X - [Best free AI with no subscription](https://benchr.org/articles/best-free-ai-no-subscription): The tools genuinely free with no credit card in 2026, the exact point where each free tier taps out, and which to pick for what. - [Best AI for writing anything long](https://benchr.org/articles/best-ai-for-writing): Opus 4.8 leads on voice; Sonnet 4.6 is the cheap pick — ranked by job: drafting, polishing, and sustained work. - [Best AI for students](https://benchr.org/articles/best-ai-for-students): NotebookLM for notes, ChatGPT 5.5 for practice problems, plus what counts as fair use versus an academic-integrity risk. - [Best AI for resumes and cover letters](https://benchr.org/articles/best-ai-for-resumes): Claude Opus 4.8 is the strongest AI for tailoring a resume to a job; Teal is the best dedicated tool, plus the ATS rules that sink an application. - [Best AI for email](https://benchr.org/articles/best-ai-for-email): When Gmail's free Gemini and Outlook's Copilot cover drafting and triage, and when a standalone like Superhuman is worth it. - [Best AI for spreadsheets and formulas](https://benchr.org/articles/best-ai-for-spreadsheets): Microsoft 365 Copilot in Excel leads; Claude runs second on big CSVs, plus where Gemini and chat models win. - [Best AI for research without the fake citations](https://benchr.org/articles/best-ai-for-research): NotebookLM, Perplexity, Elicit, Consensus, Semantic Scholar, and where each one fabricates references. - [Best AI for Arabic-English translation](https://benchr.org/articles/best-ai-for-arabic-translation): Which models move cleanly between Arabic and English both ways, and where every one of them still breaks. - [Best AI for Saudi and Gulf Arabic](https://benchr.org/articles/best-ai-for-gulf-arabic): Which AI models hold Khaleeji (Gulf) Arabic and which slide back into MSA or drift toward Egyptian. - [Best AI for customer service](https://benchr.org/articles/best-ai-for-customer-service): Off-the-shelf bots like Intercom Fin at $0.99 a resolution, platform agents, or building your own on Claude Haiku 4.5. - [Best free coding model: DeepSeek vs Qwen vs Kimi](https://benchr.org/articles/best-free-coding-model): Three open-weight models on SWE-bench Verified, license, and context, with a clear pick. - [Best AI for video in 2026](https://benchr.org/articles/best-ai-for-video): Google Veo 3.1 leads while Sora is being discontinued — Veo, Runway, Kling, Luma, and Pika with honest limits and pricing. - [Best AI tools for social media](https://benchr.org/articles/best-ai-for-social-media): Matched to platform and job: which AI tools earn a spot for captions, hooks, LinkedIn posts, and turning long video into short clips. - [Best free AI for coding](https://benchr.org/articles/best-free-ai-for-coding): What you actually get at $0 from Copilot, Cursor, Windsurf, Gemini Code Assist, and the BYO-key editors, and when the meter starts. ## Analysis - [The price-per-use-case table](https://benchr.org/articles/price-per-use-case): The cheapest model for chat, coding, RAG, agents, classification, and summarization. - [The open-weight tier right now](https://benchr.org/articles/open-weight-tier-right-now): Where Llama 4, Mistral Large 2, DeepSeek-V3.1, and Qwen 3 stand against the closed labs. - [Cutting your token bill](https://benchr.org/articles/reduce-token-usage): Where AI token spend comes from, and the levers that cut it: routing, prompt caching, the Batch API, shorter output, and lower effort. - [Why benchmarks stopped telling you anything](https://benchr.org/articles/why-benchmarks-stopped-telling-you): MMLU is saturated above 90%. The benchmarks worth tracking now. - [Small language models](https://benchr.org/articles/small-language-models): Phi-4, Gemma 3, and the documented workloads where sub-10B-parameter models can be a practical fit. - [Running models on your own machine](https://benchr.org/articles/running-models-on-your-own-machine): Hardware, software, reproducible public performance references, and when local inference may be worth evaluating. - [Renting a GPU vs. paying per token](https://benchr.org/articles/self-host-vs-api-gpu-cost): The break-even math for self-hosting an open model on a rented GPU versus a per-token API, and why utilization decides it. - [Your computer can't run the big open models](https://benchr.org/articles/run-big-models-without-a-gpu): Why DeepSeek and Llama 70B won't load on your laptop, and the four real fixes: quantize, pick a smaller model, use a hosted API, or rent a cloud GPU. - [Fine-tuning an open model without a GPU](https://benchr.org/articles/fine-tune-open-model-no-gpu): QLoRA memory arithmetic, a reproducible rented-GPU workflow, and dated editorial cost-planning scenarios rather than price guarantees. - [“CUDA out of memory”, fixed](https://benchr.org/articles/cuda-out-of-memory): The five real causes of the GPU OOM error when you serve an LLM, the fixes in order from free to last resort, and when you just need a bigger GPU. - [AI agents, eighteen months in](https://benchr.org/articles/ai-agents-eighteen-months-in): LangGraph, OpenAI Assistants v2, Anthropic computer use, and Autogen, after the hype cycle. - [RAG vs fine-tuning](https://benchr.org/articles/rag-vs-fine-tuning): When to retrieve, when to fine-tune, and the cases where fine-tuning earns its keep. - [Prompt engineering did not die](https://benchr.org/articles/prompt-engineering-did-not-die): Three prompt patterns framed as controlled A/B hypotheses, with held-out cases and measurable failure criteria rather than a universal improvement claim. - [The million-token context marketing](https://benchr.org/articles/million-token-context-marketing): What long context is actually good for, and where retrieval still beats it. - [AI for Arabic content](https://benchr.org/articles/ai-for-arabic-content): How five frontier models handle Modern Standard, Khaleeji, Egyptian, and Levantine Arabic. - [The AI agent that checks out for you](https://benchr.org/articles/agentic-shopping): How agentic shopping works end to end, who's building it, and where a hands-off purchase can go wrong. - [Are AI hallucinations fixed yet?](https://benchr.org/articles/ai-hallucinations-2026): Grounded summarization got near-perfect, but open-ended factual answers and some reasoning models still miss. - [Which AI providers train on your chats](https://benchr.org/articles/ai-privacy-who-trains-on-you): A provider-by-provider guide to which chatbots train on conversations by default and how to opt out. - [Do AI text detectors actually work?](https://benchr.org/articles/do-ai-detectors-work): Why AI detectors flag innocent students, how badly they miss non-native writers, and why a flag isn't proof. - [Do you actually need a reasoning model?](https://benchr.org/articles/do-you-need-reasoning-models): Thinking models bill hidden reasoning at the output rate and can run minutes slower — a buy-or-skip guide with real numbers. - [How to get cited inside AI answers](https://benchr.org/articles/get-cited-by-ai-search): GEO and AEO tactics that get pages quoted by ChatGPT, Perplexity, and AI Overviews, backed by the Princeton GEO study. - [When the model remembers you](https://benchr.org/articles/persistent-memory): How persistent memory works across separate chats in ChatGPT, Gemini, and Claude, and where to control it. - [What zero-click search did to the web](https://benchr.org/articles/zero-click-search): AI Overviews and chat answers now sit above the links, and the lost-click numbers are real: 58% lower CTR for the top result. ## Tools and reference - [Compare models directly](https://benchr.org/compare): Interactive comparison of pricing, benchmarks, context windows, and capability ratings for 29 frontier and open-weight models. - [OpenAI provider guide](https://benchr.org/providers/openai): Current OpenAI model IDs, list prices, context, status, reviews, and decision routes. - [Anthropic provider guide](https://benchr.org/providers/anthropic): Current Claude model IDs, list prices, context, lifecycle links, reviews, and migration routes. - [Google Gemini provider guide](https://benchr.org/providers/google-gemini): Current Gemini and Gemma identifiers, list prices, context, status, and reviews. - [Open-weight model guide](https://benchr.org/providers/open-models): License-visible open-weight records with IDs, context, hosted pricing when published, and deployment routes. - [AI model identifier finder](https://benchr.org/api-model-ids): Resolve a model name, slug, hosted API ID, or repository ID; copy the recorded value with its official source and verification date, then open pricing or the model record. A CSV export is available. - [Developer data](https://benchr.org/developers): Public read endpoints, model and lifecycle CSV exports, RSS and Atom feeds, and the deprecation calendar. - [Recent model releases](https://benchr.org/recent-releases): Major AI model launches since early 2026, in reverse chronological order. - [AI model deprecations](https://benchr.org/deprecations): Every announced model retirement across Anthropic, OpenAI, and Google — dates, replacements, and migration cost analysis, verified against official deprecation docs. - [Claude Sonnet 4 retirement](https://benchr.org/deprecations/claude-sonnet-4): Retires June 15, 2026; migration guide to Sonnet 4.6 at the same $3/$15 price. - [Claude Opus 4 and 4.1 retirement](https://benchr.org/deprecations/claude-opus-4): June 15 and August 5, 2026; Opus 4.8 replaces both at a third of the price. - [GPT-4o shutdown](https://benchr.org/deprecations/gpt-4o): The gpt-4o-2024-05-13 snapshot retires October 23, 2026; replacement options priced. - [OpenAI October 2026 retirements](https://benchr.org/deprecations/openai-october-2026-retirements): Nine model IDs retire October 23, 2026 — the end of the GPT-4 era, plus the Assistants API sunset. - [Gemini 2.5 Pro and Flash shutdown](https://benchr.org/deprecations/gemini-2-5-pro): October 16, 2026 retirement, with replacement paths that raise list prices. - [AI API price history](https://benchr.org/price-history): Append-only log of verified AI API pricing events with official sources; open data under CC BY 4.0. ## API errors - [AI API Error Database](https://benchr.org/errors): 15 common errors across OpenAI, Anthropic, and Gemini — verified causes, code fixes, and migration alternatives, filterable by provider and category. - [OpenAI insufficient_quota](https://benchr.org/errors/openai-insufficient-quota): The 429 that backoff can't fix — billing exhausted, with the code guard that separates it from rate limits. - [OpenAI model_not_found](https://benchr.org/errors/openai-model-not-found): Why model IDs 404 in 2026 — usually retirement, with the October 23 wave checklist. - [Anthropic overloaded_error 529](https://benchr.org/errors/anthropic-overloaded-error): Platform-wide overload vs your account, and how to retry without making it worse. - [Anthropic invalid_request_error](https://benchr.org/errors/anthropic-invalid-request-error): The 2026 causes — sampling params on Opus 4.7+, prefill, modified thinking blocks. - [Gemini RESOURCE_EXHAUSTED](https://benchr.org/errors/gemini-resource-exhausted): Free-tier rate limits and the quota decision tree. ## About - [About benchr](https://benchr.org/about): What benchr is and how it sources information. - [benchr Editorial Team](https://benchr.org/editorial-team): The accountable publication-level author — roles, process, AI-assistance disclosure, and correction routes. - [Contact](https://benchr.org/contact): General questions, corrections with priority handling, and security reports. - [Methodology](https://benchr.org/methodology): Where the data comes from and how it is kept current. - [Editorial standards](https://benchr.org/editorial-standards): Publishing principles and sourcing rules. - [Corrections](https://benchr.org/corrections): The log of material corrections. - [Affiliate disclosure](https://benchr.org/affiliate-disclosure): Revenue relationships, the partners currently used, and the firewall between compensation and editorial rankings.