benchr

What people are doing with AI

Things somebody built, ran or published, each linked back to whoever did it. benchr has not reproduced any of it, and every record says so.

Topic

What happened

57 working web demos, built by agent fleets in one day

Reported by elder-pliniuslaunch week of September 1, 2026Not tested by benchr

A public showcase collects 57 self-contained web demos across six collections: generative art, simulations, instruments and explorable worlds. The page states each piece was designed, built and QA'd end to end by fleets of Claude Fable 5 agents, over one day, and that every demo is a single HTML file with no dependencies.

Why it matters The review step was handled by the agents too, which is the part that usually does not survive contact with a deadline.

Original source: elder-plinius.github.io

How it was done

  1. InSix themes: generative art, simulations, instruments, worlds
  2. ThenFleets of agents designed, built and reviewed each piece
  3. ConstraintOne HTML file per demo, no dependencies
  4. Out57 demos, published in one day

From one photo of an empty lot to a filmed tour of a house

Reported by @alexalbert__launch week of September 1, 2026Not tested by benchr

Alex Albert gave Claude Fable 5.1 a single photograph of an empty property lot. The model designed a house to fit the lot, rendered it, and produced a cinematic walkthrough video of the result. The work was done through code rather than in a design tool.

  1. InOne photograph of an empty lot
  2. ThenThe model wrote code instead of opening a design tool
  3. ThenIt designed a house to fit the lot and rendered it
  4. OutA cinematic walkthrough video

An agent you talk to on WhatsApp is now one of the most-starred projects on GitHub

Reported by Peter SteinbergerAugust 30, 2026Not tested by benchr

OpenClaw is a free, open-source autonomous agent whose interface is a messaging app: you talk to it on Signal, Telegram, Discord or WhatsApp and it runs tasks through whichever model you point it at. Austrian programmer Peter Steinberger released it on November 24, 2025 as Warelay, renamed it twice, and shipped version 2.0 on August 30, 2026. benchr read the repository API directly on 2026-09-04: 388,822 stars and 81,662 forks, created 2025-11-24 — the date the encyclopaedia entry gives independently.

  1. InA message in Signal, Telegram, Discord or WhatsApp
  2. ThenA skills system decides what to run
  3. ThenThe model you configured does the work; config stays local
  4. OutThe reply comes back in the same chat

More of what people did

On the hardest web tasks, people finish 10%. The best agent finishes 8%.

Reported by Zejun Xu and colleaguesAugust 9, 2026Not tested by benchr

CAP, published on arXiv on August 9, 2026, is a benchmark of 420 cross-site tasks on 108 real websites across 24 domains, built around interactions that need real visual understanding rather than form-filling. Full success rates: Manus 8.0%, Fellou 7.0%, Comet 6.0%, Genspark 5.0%, Claude-4.5-Sonnet 5.0%, Dia 4.0%, DeepSeek-V4-Flash 2.9%, GPT-5 2.0% — against a human baseline of 10.0%. On partial completion the gap opens up: humans 35.0%, Comet 48.0%, Manus 23.0%.

  1. In420 tasks, 108 real websites, 24 domains
  2. ThenSplit into complex actions and complex perception
  3. ThenEight agents and a human baseline run against the same set
  4. OutPerception, not manipulation, is the dominant failure

16 working viruses, designed by a genome language model

Reported by Samuel King and Brian Hie (Stanford / Arc Institute)August 6, 2026Not tested by benchr

A team at Stanford and the Arc Institute used the genome language models Evo 1 and Evo 2 to generate roughly 700,000 candidate bacteriophage genomes. They selected 302 for synthesis, successfully built 285, and introduced them into E. coli. Sixteen produced viable phages. Combined into one cocktail, those sixteen cleared two E. coli strains that had already become resistant to a naturally occurring phage. The training data deliberately excluded viruses able to infect humans, animals or plants.

  1. InGenome language models Evo 1 and Evo 2
  2. Then~700,000 candidate phage genomes generated
  3. Then302 selected, 285 successfully synthesised
  4. Out16 produced viable phages, which cleared resistant E. coli

Agents spend 10 million tokens and 85 minutes on one task. The best finishes 15% of them.

Reported by Zongxia Li and colleaguesJuly 9, 2026Not tested by benchr

Long-Horizon-Terminal-Bench, published on arXiv on July 9, 2026, grades agents on 46 terminal tasks across nine categories with partial credit rather than pass or fail. Fifteen frontier models were run. Against a 0.95 reward threshold the strongest model passes 15.2% of tasks first time, and 10.9% against a perfect score. The mean across all fifteen is 4.3% and 1.7%. A single task consumes on average 9.9 million tokens, about 231 episodes and 85.3 minutes of execution.

  1. In46 terminal tasks across nine categories
  2. ThenGraded with dense partial credit, not pass or fail
  3. Then15 frontier models, ~231 episodes and 85 minutes per task
  4. OutBest 15.2% pass@1; mean across models 4.3%

A 95-minute feature film, 15 people, two weeks

Reported by Aitore ZholdaskaliMay 16, 2026Not tested by benchr

Hell Grind, an action fantasy directed by Aitore Zholdaskali and co-written with Adilkhan Yerzhanov, was made by a crew of fifteen using Higgsfield's Soul Cinema and Soul Cast alongside the Dreamina Seedance 2.0 video generator. The budget was $500,000, of which $400,000 went on AI compute. Production took two weeks. The finished film runs 95 minutes and was screened for the industry in Cannes on May 16, 2026, during the festival but not part of it.

  1. InA prompt per shot, averaging 3,000 words
  2. ThenEach prompt specified cinematography and physics constraints
  3. ThenThe generator returned 15-second clips, run again and again
  4. ThenA human picked the takes — the CEO called it "a feeling of a slot machine"
  5. Out95 minutes, cut from two weeks of takes

They gave a coding agent the eval harness and let it tune the agent

Reported by Magnus Müller, Browser UseMarch 25, 2026Not tested by benchr

Magnus Müller of Browser Use describes handing Claude Code a command-line interface to their own evaluation platform plus a prompt to run in a loop, twenty cycles, in parallel — "a search tree over the space of possible best agents". The resulting agent scored 97% on Online-Mind2Web, a benchmark of 300 tasks across 136 real websites (83 easy, 143 medium, 74 hard). The post calls it the highest score ever recorded on it.

  1. InA browser agent that already worked, and its eval platform
  2. ThenClaude Code got a CLI to the evals and a prompt to run in a loop
  3. ThenTwenty cycles, in parallel, split into train and validation sets
  4. Out97% on Online-Mind2Web

Also this month

When it does not work

The failures people actually hit, what causes them, and what to do instead.

See all

Moves you can copy

Each one names what goes in, what the move is, and what comes out different.

See all

Claims we checked

MISLEADING

The AI-designed bacteriophages were genuinely new organisms, carrying genes nothing in nature had written.

What is actually true

The genomes were generated by a model and 16 of them produced working phages — that part holds. But an independent analysis by Oliver Crook at Oxford found the 16 viable genomes averaged about 97 percent identity to their ΦX174 template, placing them inside existing phage diversity rather than outside it. Crook's summary was that these were "brothers and sisters of the original virus". The genuinely novel result was recombination, not invention: one design took an unusually truncated protein from a distant phage species and made it work on the ΦX174 backbone, which conventional genetic engineering had failed to do.

Where it spread

  • Popular science coverage and aggregator summaries of the August 2026 Science paper
  • Social posts framing the result as the first artificial life written by AI

Checked

September 3, 2026 · Science · IEEE Spectrum