How verification works here

Two questions, kept apart on purpose: does the provider document it, and has benchr run it. The second is the harder one, and benchr has not answered it yet for any capability.

These records describe what providers officially document. benchr has not run these capabilities itself, so nothing here is a test result. Each record shows the vendor page it was read from and the date.

Ledger updated: September 3, 2026

What counts as evidence here

Six different things get called “checked”. Reading a provider's page is not running the capability, and fetching a URL is not reproducing a result. Each run below says which one it is.

benchr ran it
benchr ran the capability itself and recorded what came back. Only this kind can earn Benchr Verified.
benchmark run
benchr executed a published benchmark under a stated protocol and reported the score it got, not the score somebody else published.
regression check
A capability that worked before was run again to see whether it still does. The value is in the comparison, not the single result.
claim checked
A specific factual claim benchr publishes was tested against the world - a URL fetched, a count read from an API, an artefact opened. It settles a fact; it never promotes a capability.
documentation read
A provider's own page was read and what it says was recorded, with the URL and the date. This is reading, not testing, and benchr says so wherever it appears.
read by a person
A person looked and formed a judgement. Useful, and the weakest thing here: it does not reproduce, and another person might see it differently.

Where verification stands today

No verification run has been accepted yet, so no capability carries the Benchr Verified badge. All 20 records sit at “Not tested by benchr” on the testing axis.

7 runs below are self-reported: they were produced by the same automated session that maintains this repository, which makes them a starting point and not evidence. They are published anyway, in full, because a held run nobody can read is indistinguishable from no run at all.

Designed and not run

These are written out in full before they can be run, so nothing about the test can be decided once the result is known. None of them is evidence, and none of them can promote anything.

The benchr verification protocol

Seven steps. A run is void if any of them is missing from the record.

  1. State the claim as a single testable sentence before running anything, so the run cannot be reinterpreted after the fact.
  2. Name the subject exactly: model and version, the surface or interface, and the version of that surface.
  3. Record the environment: operating system, network conditions, and anything that could make the result local rather than general.
  4. Record the inputs and the exact instruction or prompt, verbatim. A paraphrased prompt is not evidence.
  5. Record the settings that were in force, including any left at their defaults.
  6. Record the observed output and the result. Partial and fail are published exactly like a pass.
  7. Record the evidence and the reproduction recipe, then list the limitations of the run - starting with what it did not test.

Documentation axis

  • DocsOfficially documentedThe provider's own current documentation describes this. benchr read that page on the date shown.
  • DocsDocumented in partThe provider documents the mechanism but not the result, or documents it only inside stated limits.
  • DocsDeprecatedThe provider has announced withdrawal, closed new access, or set a retirement date.
  • DocsNeeds recheckThe documentation has not been re-read inside the recheck window. Computed, never hand-set.
  • DocsNot documentedNo official source describes this. Submissions land here until one is found.

Testing axis

  • TestBenchr Verifiedbenchr ran this and reproduced the result. Every run is published with its model, surface, inputs, settings, environment, output and evidence.
  • TestPartly reproducedA benchr run succeeded under narrower conditions than the claim.
  • TestNeeds retestThe last passing run is older than the retest window, or a dependency changed.
  • TestBrokenA benchr run failed. The run is published alongside the record.
  • TestNot tested by benchrbenchr has not run this. This is the honest default and it is where every record starts.