benchr compiles reference information on AI language models — pricing, context windows, benchmark scores, and deprecation status. Official prices and specifications come from provider documentation, third-party benchmark results come from the benchmark publisher or maintainer, and benchr's own estimates are identified as editorial estimates. Coverage includes Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen, Phi, and other frontier and open-weight models.
benchr is a working reference, not a benchmarking lab. Each article separates documented capabilities, published prices, and known limits from editorial analysis. If a page includes original evaluation, it also explains the test setup.
benchr also publishes original evaluation materials: six bilingual suites containing fixed prompts and scoring rubrics. The public package contains no model outputs, scores, or rankings, and publishing a reusable test design is not a claim that benchr ran the models.
Articles are currently credited to the benchr Editorial Team, a publication-level Organization with documented editorial roles and a public correction route. The byline does not invent a person, credential, or staff size.
What you'll find
You'll find provider-sourced pricing tables, context-window comparisons, release histories, and deprecation notes. Pricing links back to pages such as Anthropic's pricing page, OpenAI's API pricing, and Google's Gemini API models page.
Benchmark figures point to their publisher or maintainer, including SWE-bench Verified, LMSYS Arena, and ARC-AGI. Use-case recommendations explain the evidence and trade-offs behind the shortlist.
The capability record
Alongside the model data, benchr keeps a record of what AI can actually be made to do: 20 documented capabilities, 23 goal-first tasks, 9 multi-step workflows, and 7 tools and surfaces they run on. Around that record sit 12 discoveries — things other people did, each linked back to whoever did it — 16 problems with what to do about them, and 12 techniques written as input, method and result.
benchr has not tested any of these capabilities. A discovery, a problem and a technique are not benchr test results either: a discovery is somebody else’s claim forwarded with attribution, and a problem or technique says on its face whether its cause is documented by a vendor or is a pattern nobody documents. Every record on those pages describes what a provider officially documents, read from the vendor's own page on the date shown — that is what Officially documented means, and it is the strongest claim the ledger currently makes. Where a claim cannot be sourced, the record lists it under "not stated by the source" instead of filling the gap, and a record whose documentation has not been re-read inside the recheck window is marked Needs recheck automatically rather than by anyone's decision. When a provider withdraws something — as OpenAI did with new fine-tuning access — the record stays, marked Withdrawn, because the history is the point. A separate first-party testing programme is being built; until a test is published and reproducible, no page will say benchr ran anything. Anyone can submit something for review.
What you won't find
Synthetic precision claims. Invented test costs. Undisclosed paid placement. Rankings sold to vendors. Unlabelled affiliate links. Empty doorway pages built to catch a search query without answering it.
Use-case guides answer practical questions such as which models to test for writing, coding, or research. Each guide needs its own evidence, prices, trade-offs, and test plan. A page does not belong here if it only swaps product names into a template.
Independence and funding
benchr is an independent editorial publication. Some clearly marked links to AI/ML API, Lovable, and RunPod are affiliate links; benchr may earn a commission when a reader uses one, at no extra cost to the reader. Those relationships do not buy coverage, rankings, favorable conclusions, or review rights. Google AdSense is a planned source of revenue after approval, but AdSense ads are not currently loaded. See the full affiliate disclosure.
Use of AI
AI tools may assist with outlining, draft language, translation checks, code, and custom illustrations. They are not authors, sources, or independent reviewers. Factual claims such as prices, benchmark results, and dates are checked against the cited provider or benchmark source before publication, editorial estimates are identified as estimates, and responsibility remains with the editorial team.
Authorship and editorial responsibility
The editorial-team name represents responsibilities for scope, evidence checks, Arabic editing, release checks, and corrections; it does not claim that each role is a separate employee. A real Person byline will be used only when the owner supplies a public name, role, profile URL, and explicit article assignment. Until then, the Organization byline is the truthful public author identity.
Publication identity
benchr is the public name of this independent publication. General questions go to hello@benchr.org, corrections to corrections@benchr.org, and security reports to security@benchr.org. Where Google or another platform requires legal operator details, the account and payments profile must use the operator's exact legal name and address and match the submitted identity, tax, address, and payment records; the benchr brand does not replace those legal records.
Article datelines distinguish two things: "Updated" means the content itself changed, logged in that article's changelog, while "Reviewed" means the facts were re-checked against primary sources on that date without the content changing. A site-wide verification pass ran on May 30, 2026.
Corrections
Errors are corrected on the original article with a note in the article's changelog. Material corrections also appear on the corrections page.
Contact
For corrections or source disputes: corrections@benchr.org.