One screenshot in, a working page out
Hand a model a picture of a screen and get markup and styles back that reproduce the layout, spacing and copy closely enough to edit.
Each record is one concrete thing a provider documents a model or agent can do — with the steps, the limits, the surfaces it runs on, and the official page benchr read it from. None of it has been tested by benchr yet.
These records describe what providers officially document. benchr has not run these capabilities itself, so nothing here is a test result. Each record shows the vendor page it was read from and the date.
Ledger updated: September 3, 2026
Hand a model a picture of a screen and get markup and styles back that reproduce the layout, spacing and copy closely enough to edit.
The model gets a live browser: it navigates, reads the accessibility tree, clicks by element reference, fills forms and manages tabs, in a loop, until the task is done.
Screenshot, move, click, type. The model works a native application the way a person does, for software that never got an API.
Supply a schema and the model's response adheres to it - no missing required key, no invented enum value, no defensive parser.
Instead of describing an analysis, the model executes it in a sandbox, reads the actual output, and corrects itself before answering.
Not OCR. The model reads the text and looks at each page, so a figure that only exists as a chart is still answerable.
The model decides when to search, the API runs the searches, and the final answer carries source URLs, titles and the exact quoted span.
One open protocol between an AI application and your systems, so a connector you write once works in every client that speaks it.
Mark the stable part of a prompt as cacheable and every later request reads it at 0.1x the input price instead of paying full price again.
Submit many requests as one asynchronous batch, poll for completion, and pay 50% less than the same work sent one request at a time.
An agent that reads the codebase itself, edits across many files, runs the commands, reads the failures, and opens the pull request.
A folder with a SKILL.md, optional reference files and scripts. The model reads the name and description always, the instructions only when your request matches.
Speech to text is the easy half. The useful half is a transcript where each line is attributed to a speaker.
Not transcribe-then-answer-then-synthesise. One session on gpt-realtime-2.1 where audio goes in and audio comes back, with the turn-taking handled for you.
The failure everyone remembers from early image models - garbled lettering - is documented as a supported capability now, for infographics, menus, diagrams and marketing assets.
Text or an image goes in, a short video comes out - and on one of the two documented models, with native audio.
A whole codebase, a full deposition, an hour of video. The window is real; the retrieval behavior inside it is the part people get wrong.
Open-weight models on your own hardware, where the constraint is memory arithmetic rather than a price per token.
Start the work in the terminal, move it to the cloud, check it from a phone, and have it run again next Tuesday without you.
Closed to new users. The documentation states the fine-tuning platform is being wound down and is no longer accessible to new users, with existing users able to create jobs temporarily.
Nothing matches that yet
The ledger holds 20 documented capabilities, so a narrow filter empties quickly. That is the honest state, not a search failure.