<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>benchr</title>
    <link>https://benchr.org/</link>
    <description>Pricing, benchmarks, and use-case fit for AI model selection.</description>
    <language>en-us</language>
    <lastBuildDate>Mon, 07 Sep 2026 21:14:01 +0000</lastBuildDate>
    <atom:link href="https://benchr.org/feed.xml" rel="self" type="application/rss+xml" />
    <managingEditor>corrections@benchr.org (benchr)</managingEditor>
    <item>
      <title>Kimi K2.7-Code: fewer thinking tokens, and a flat 2x for speed</title>
      <link>https://benchr.org/articles/kimi-k2-7-code-review</link>
      <guid isPermaLink="true">https://benchr.org/articles/kimi-k2-7-code-review</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate>
      <description>Moonshot&#x27;s coding model claims ~30% fewer thinking tokens than K2.6, on a rate card where output costs 4.2x input. The high-speed variant doubles every line.</description>
    </item>
    <item>
      <title>Gemini 3.8 Flash: same rate card, different bill</title>
      <link>https://benchr.org/articles/gemini-3-8-flash-review</link>
      <guid isPermaLink="true">https://benchr.org/articles/gemini-3-8-flash-review</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate>
      <description>Gemini 3.8 Flash is priced to the cent like 3.7 Flash. Google says it spends more tokens on purpose, which is where the cost difference lives.</description>
    </item>
    <item>
      <title>DeepSeek-V4-Flash-Vision-Exp: vision at no premium</title>
      <link>https://benchr.org/articles/deepseek-v4-flash-vision-exp-review</link>
      <guid isPermaLink="true">https://benchr.org/articles/deepseek-v4-flash-vision-exp-review</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate>
      <description>DeepSeek charges nothing extra for image input on this checkpoint. What you trade is stability: it is experimental and has no published license.</description>
    </item>
    <item>
      <title>Write the shot, not the vibe</title>
      <link>https://benchr.org/techniques/write-the-shot-not-the-vibe</link>
      <guid isPermaLink="true">https://benchr.org/techniques/write-the-shot-not-the-vibe</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>A shot description that fixes camera, lens, movement, light and the constants is doing the job a shot list and a cinematographer would do. Short prompts are…</description>
    </item>
    <item>
      <title>Write the repo rules down once</title>
      <link>https://benchr.org/techniques/write-the-repo-rules-down-once</link>
      <guid isPermaLink="true">https://benchr.org/techniques/write-the-repo-rules-down-once</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Put the build command, the test command, the untouchable files and the definition of a complete change in a file at the repository root. Agents read it. So…</description>
    </item>
    <item>
      <title>They gave a coding agent the eval harness and let it tune the agent</title>
      <link>https://benchr.org/discover/an-agent-tuned-another-agent-until-it-stopped-failing</link>
      <guid isPermaLink="true">https://benchr.org/discover/an-agent-tuned-another-agent-until-it-stopped-failing</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>The optimiser was not a training run or a hyper-parameter sweep. It was a coding agent editing another agent&#x27;s prompts and tools, judged by a benchmark.</description>
    </item>
    <item>
      <title>The video shimmers — textures crawl and the light flickers</title>
      <link>https://benchr.org/fix/the-video-shimmers-and-nothing-holds-still</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-video-shimmers-and-nothing-holds-still</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Shorten the shot. Most shimmer becomes invisible below about two seconds, and a cut costs nothing.</description>
    </item>
    <item>
      <title>The prototype cost pennies and the real thing costs a fortune</title>
      <link>https://benchr.org/fix/the-bill-was-much-bigger-than-expected</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-bill-was-much-bigger-than-expected</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Find the largest unchanging block in your prompt. If it is re-sent every turn and not cached, that single change is usually most of the bill.</description>
    </item>
    <item>
      <title>The product looks different in every shot of my ad</title>
      <link>https://benchr.org/fix/the-video-changes-the-thing-between-shots</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-video-changes-the-thing-between-shots</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Generate from a still of the actual product rather than from a description of it, for every single shot.</description>
    </item>
    <item>
      <title>The poster looks great until you read the words on it</title>
      <link>https://benchr.org/fix/the-text-inside-the-image-is-gibberish</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-text-inside-the-image-is-gibberish</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Keep the text to a few words, put it in straight quotation marks in the prompt, and say where on the image it goes.</description>
    </item>
    <item>
      <title>The eyes change color and the kid looks older on every page</title>
      <link>https://benchr.org/fix/the-eyes-and-hair-drift-across-a-set</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-eyes-and-hair-drift-across-a-set</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Name the attributes that must not move — eye color, hairline, one distinctive mark — and check those three specifically rather than judging the face as a whole.</description>
    </item>
    <item>
      <title>The browser agent clicks the wrong button</title>
      <link>https://benchr.org/fix/the-agent-clicks-the-wrong-thing</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-agent-clicks-the-wrong-thing</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Refer to targets by their accessible name, not by where they are. &#x27;The button labeled Continue&#x27; survives a layout change; &#x27;the button on the right&#x27; does not.</description>
    </item>
    <item>
      <title>Schema first, prompt second</title>
      <link>https://benchr.org/techniques/schema-first-prompt-second</link>
      <guid isPermaLink="true">https://benchr.org/techniques/schema-first-prompt-second</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Write the output schema before the instruction. The prompt then only has to describe the judgement, because the shape is already enforced.</description>
    </item>
    <item>
      <title>Same character, second image, different face</title>
      <link>https://benchr.org/fix/the-person-changes-between-images</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-person-changes-between-images</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Stop re-describing the person and start re-supplying them: pass the first image back in as a reference for every subsequent generation.</description>
    </item>
    <item>
      <title>Quote the string, or leave the space</title>
      <link>https://benchr.org/techniques/quote-the-string-leave-the-space</link>
      <guid isPermaLink="true">https://benchr.org/techniques/quote-the-string-leave-the-space</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Two moves, in order of reliability: quote a short string exactly, or generate the composition with deliberate empty space and set the type in an editor.</description>
    </item>
    <item>
      <title>Put the rule last</title>
      <link>https://benchr.org/techniques/put-the-rule-last</link>
      <guid isPermaLink="true">https://benchr.org/techniques/put-the-rule-last</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Instructions placed after a long attachment are followed more consistently than the same instructions placed before it. Recency is doing work that emphasis…</description>
    </item>
    <item>
      <title>Outline first, then one unit per turn</title>
      <link>https://benchr.org/techniques/outline-then-one-unit-per-turn</link>
      <guid isPermaLink="true">https://benchr.org/techniques/outline-then-one-unit-per-turn</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Split the job before the output limit splits it for you, at a boundary you chose rather than mid-sentence.</description>
    </item>
    <item>
      <title>Open-source projects are now writing their contributor guides for agents</title>
      <link>https://benchr.org/discover/maintainers-started-writing-the-rules-for-robots</link>
      <guid isPermaLink="true">https://benchr.org/discover/maintainers-started-writing-the-rules-for-robots</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>The quote that explains the whole shift is Tindle&#x27;s: an agent-written pull request is &quot;basically somebody else paying for your compute&quot;.</description>
    </item>
    <item>
      <title>One reference, one variable</title>
      <link>https://benchr.org/techniques/one-reference-one-variable</link>
      <guid isPermaLink="true">https://benchr.org/techniques/one-reference-one-variable</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Stop describing the subject and start re-supplying it. Then change exactly one thing per generation so you can tell what caused a drift.</description>
    </item>
    <item>
      <title>On the hardest web tasks, people finish 10%. The best agent finishes 8%.</title>
      <link>https://benchr.org/discover/the-web-tasks-that-stop-agents-stop-people-too</link>
      <guid isPermaLink="true">https://benchr.org/discover/the-web-tasks-that-stop-agents-stop-people-too</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>One agent beats the human baseline on partial progress and still loses on finishing.</description>
    </item>
    <item>
      <title>My coding agent fixed the bug and broke three other things</title>
      <link>https://benchr.org/fix/the-agent-broke-code-that-was-working</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-agent-broke-code-that-was-working</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Work on a branch, and read the diff before the tests. A large diff for a small request is the signal, not the test result.</description>
    </item>
    <item>
      <title>Make it restate the brief</title>
      <link>https://benchr.org/techniques/make-it-restate-the-brief</link>
      <guid isPermaLink="true">https://benchr.org/techniques/make-it-restate-the-brief</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>One paragraph, in its own words, before it starts. You find out whether the constraints landed while it is still cheap to correct.</description>
    </item>
    <item>
      <title>It stops mid-task and tells me I have hit a limit</title>
      <link>https://benchr.org/fix/the-run-dies-halfway-because-i-hit-a-limit</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-run-dies-halfway-because-i-hit-a-limit</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Read the error before retrying. A 429 with a retry hint wants a wait; a quota error wants a smaller job, because no amount of backoff refills a balance.</description>
    </item>
    <item>
      <title>It scores brilliantly on the demo and falls over on my actual work</title>
      <link>https://benchr.org/fix/the-agent-passes-the-test-and-fails-the-job</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-agent-passes-the-test-and-fails-the-job</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Take ten tasks from your own week, write down what a correct answer looks like, and run those. That is your benchmark; everyone else&#x27;s measures someone…</description>
    </item>
    <item>
      <title>It says it read my PDF, but it only answered about the first few pages</title>
      <link>https://benchr.org/fix/the-pdf-comes-back-half-read</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-pdf-comes-back-half-read</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Ask it one question it can only answer from the end of the document — a figure from the last table, the final section heading. If it cannot, the file did…</description>
    </item>
    <item>
      <title>It returns JSON most of the time, then wraps it in prose and breaks my parser</title>
      <link>https://benchr.org/fix/the-json-comes-back-broken</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-json-comes-back-broken</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Stop asking for JSON and start supplying a schema. The API feature is called structured outputs and it is a request parameter, not a prompt trick.</description>
    </item>
    <item>
      <title>It gets most of the way through, then just stops mid-sentence</title>
      <link>https://benchr.org/fix/it-stops-halfway-through-a-long-job</link>
      <guid isPermaLink="true">https://benchr.org/fix/it-stops-halfway-through-a-long-job</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Ask for one section, not the whole thing, and then ask for the next. Two short answers beat one truncated one.</description>
    </item>
    <item>
      <title>It described the chart correctly and then got the number wrong</title>
      <link>https://benchr.org/fix/it-cannot-read-the-number-off-the-chart</link>
      <guid isPermaLink="true">https://benchr.org/fix/it-cannot-read-the-number-off-the-chart</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Give it the numbers, not a picture of the numbers. A CSV export, the underlying table, or the data endpoint removes the entire failure mode.</description>
    </item>
    <item>
      <title>Identical prompt, different provider, completely different behavior</title>
      <link>https://benchr.org/fix/the-prompt-works-on-one-model-and-fails-on-another</link>
      <guid isPermaLink="true">https://benchr.org/fix/the-prompt-works-on-one-model-and-fails-on-another</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Before blaming the prompt, check that the feature it depends on exists on the target model at all. Half of these failures are a missing parameter, not a…</description>
    </item>
    <item>
      <title>Hold back a set the agent never sees</title>
      <link>https://benchr.org/techniques/hold-back-a-validation-set-the-agent-never-sees</link>
      <guid isPermaLink="true">https://benchr.org/techniques/hold-back-a-validation-set-the-agent-never-sees</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Your own tasks, split in two. You iterate against one half; the other half is only ever run, never tuned against. The second number is the real one.</description>
    </item>
    <item>
      <title>Halfway through, it stopped following the instructions I gave at the start</title>
      <link>https://benchr.org/fix/it-forgot-what-i-told-it-earlier</link>
      <guid isPermaLink="true">https://benchr.org/fix/it-forgot-what-i-told-it-earlier</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Re-state the two or three rules that matter immediately before the request, not at the top of the conversation. Recency beats position.</description>
    </item>
    <item>
      <title>Give it the data, not a picture of the data</title>
      <link>https://benchr.org/techniques/give-it-the-data-not-a-picture-of-the-data</link>
      <guid isPermaLink="true">https://benchr.org/techniques/give-it-the-data-not-a-picture-of-the-data</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>The single highest-value substitution in analysis work. A CSV removes an entire measured failure mode that no prompt fixes.</description>
    </item>
    <item>
      <title>Cache the part that never changes</title>
      <link>https://benchr.org/techniques/cache-the-part-that-never-changes</link>
      <guid isPermaLink="true">https://benchr.org/techniques/cache-the-part-that-never-changes</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Find the largest block of your prompt that is identical on every call — the system prompt, the style guide, the reference document — and stop paying full…</description>
    </item>
    <item>
      <title>Ask for the last page first</title>
      <link>https://benchr.org/techniques/ask-for-the-last-page-first</link>
      <guid isPermaLink="true">https://benchr.org/techniques/ask-for-the-last-page-first</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>A thirty-second check that tells you whether the model actually has the whole document before you trust a summary of it.</description>
    </item>
    <item>
      <title>An agent you talk to on WhatsApp is now one of the most-starred projects on GitHub</title>
      <link>https://benchr.org/discover/openclaw-outgrew-the-encyclopaedia-entry-about-it</link>
      <guid isPermaLink="true">https://benchr.org/discover/openclaw-outgrew-the-encyclopaedia-entry-about-it</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>The interface is the surprise. No terminal, no dashboard, no browser tab — the agent is a contact in a chat app you already have open.</description>
    </item>
    <item>
      <title>Agents spend 10 million tokens and 85 minutes on one task. The best finishes 15% of them.</title>
      <link>https://benchr.org/discover/ten-million-tokens-and-eighty-five-minutes-to-fail</link>
      <guid isPermaLink="true">https://benchr.org/discover/ten-million-tokens-and-eighty-five-minutes-to-fail</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>The cost is the headline, not the score. Eighty-five minutes and ten million tokens is a real bill, and it buys a first-time finish about one time in seven.</description>
    </item>
    <item>
      <title>A 95-minute feature film, 15 people, two weeks</title>
      <link>https://benchr.org/discover/a-95-minute-feature-shot-by-fifteen-people-in-two-weeks</link>
      <guid isPermaLink="true">https://benchr.org/discover/a-95-minute-feature-shot-by-fifteen-people-in-two-weeks</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>Eighty percent of the budget was compute, not people — the opposite of how a film has ever been costed.</description>
    </item>
    <item>
      <title>&quot;The CGI would have cost millions. I spent $2,000.&quot;</title>
      <link>https://benchr.org/discover/a-two-thousand-dollar-film-premiered-at-tribeca</link>
      <guid isPermaLink="true">https://benchr.org/discover/a-two-thousand-dollar-film-premiered-at-tribeca</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate>
      <description>A festival premiere is the bar that separates a demo from a film, and this one was cleared at a hobby budget.</description>
    </item>
    <item>
      <title>No publishing API, so the model built its own browser extension</title>
      <link>https://benchr.org/discover/no-api-to-publish-so-the-model-built-a-browser-extension</link>
      <guid isPermaLink="true">https://benchr.org/discover/no-api-to-publish-so-the-model-built-a-browser-extension</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate>
      <description>The missing interface did not stop the task; the model built the interface.</description>
    </item>
    <item>
      <title>Given a broken setup, the model installed its own tooling</title>
      <link>https://benchr.org/discover/the-model-installed-its-own-blender-tooling</link>
      <guid isPermaLink="true">https://benchr.org/discover/the-model-installed-its-own-blender-tooling</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate>
      <description>Repairing its own working environment was not the task it was given; it was what the task required first.</description>
    </item>
  </channel>
</rss>
