Reported by elder-plinius·launch week of September 1, 2026·Not tested by benchr
A public showcase collects 57 self-contained web demos across six collections: generative art, simulations, instruments and explorable worlds. The page states each piece was designed, built and QA'd end to end by fleets of Claude Fable 5 agents, over one day, and that every demo is a single HTML file with no dependencies.
Why it matters The review step was handled by the agents too, which is the part that usually does not survive contact with a deadline.
Reported by @alexalbert__·launch week of September 1, 2026·Not tested by benchr
Alex Albert gave Claude Fable 5.1 a single photograph of an empty property lot. The model designed a house to fit the lot, rendered it, and produced a cinematic walkthrough video of the result. The work was done through code rather than in a design tool.
InOne photograph of an empty lot
ThenThe model wrote code instead of opening a design tool
ThenIt designed a house to fit the lot and rendered it
Reported by Peter Steinberger·August 30, 2026·Not tested by benchr
OpenClaw is a free, open-source autonomous agent whose interface is a messaging app: you talk to it on Signal, Telegram, Discord or WhatsApp and it runs tasks through whichever model you point it at. Austrian programmer Peter Steinberger released it on November 24, 2025 as Warelay, renamed it twice, and shipped version 2.0 on August 30, 2026. benchr read the repository API directly on 2026-09-04: 388,822 stars and 81,662 forks, created 2025-11-24 — the date the encyclopaedia entry gives independently.
InA message in Signal, Telegram, Discord or WhatsApp
ThenA skills system decides what to run
ThenThe model you configured does the work; config stays local
Reported by Zejun Xu and colleagues·August 9, 2026·Not tested by benchr
CAP, published on arXiv on August 9, 2026, is a benchmark of 420 cross-site tasks on 108 real websites across 24 domains, built around interactions that need real visual understanding rather than form-filling. Full success rates: Manus 8.0%, Fellou 7.0%, Comet 6.0%, Genspark 5.0%, Claude-4.5-Sonnet 5.0%, Dia 4.0%, DeepSeek-V4-Flash 2.9%, GPT-5 2.0% — against a human baseline of 10.0%. On partial completion the gap opens up: humans 35.0%, Comet 48.0%, Manus 23.0%.
In420 tasks, 108 real websites, 24 domains
ThenSplit into complex actions and complex perception
ThenEight agents and a human baseline run against the same set
OutPerception, not manipulation, is the dominant failure
Reported by Samuel King and Brian Hie (Stanford / Arc Institute)·August 6, 2026·Not tested by benchr
A team at Stanford and the Arc Institute used the genome language models Evo 1 and Evo 2 to generate roughly 700,000 candidate bacteriophage genomes. They selected 302 for synthesis, successfully built 285, and introduced them into E. coli. Sixteen produced viable phages. Combined into one cocktail, those sixteen cleared two E. coli strains that had already become resistant to a naturally occurring phage. The training data deliberately excluded viruses able to infect humans, animals or plants.
InGenome language models Evo 1 and Evo 2
Then~700,000 candidate phage genomes generated
Then302 selected, 285 successfully synthesised
Out16 produced viable phages, which cleared resistant E. coli
Reported by Zongxia Li and colleagues·July 9, 2026·Not tested by benchr
Long-Horizon-Terminal-Bench, published on arXiv on July 9, 2026, grades agents on 46 terminal tasks across nine categories with partial credit rather than pass or fail. Fifteen frontier models were run. Against a 0.95 reward threshold the strongest model passes 15.2% of tasks first time, and 10.9% against a perfect score. The mean across all fifteen is 4.3% and 1.7%. A single task consumes on average 9.9 million tokens, about 231 episodes and 85.3 minutes of execution.
In46 terminal tasks across nine categories
ThenGraded with dense partial credit, not pass or fail
Then15 frontier models, ~231 episodes and 85 minutes per task
Reported by Aitore Zholdaskali·May 16, 2026·Not tested by benchr
Hell Grind, an action fantasy directed by Aitore Zholdaskali and co-written with Adilkhan Yerzhanov, was made by a crew of fifteen using Higgsfield's Soul Cinema and Soul Cast alongside the Dreamina Seedance 2.0 video generator. The budget was $500,000, of which $400,000 went on AI compute. Production took two weeks. The finished film runs 95 minutes and was screened for the industry in Cannes on May 16, 2026, during the festival but not part of it.
InA prompt per shot, averaging 3,000 words
ThenEach prompt specified cinematography and physics constraints
ThenThe generator returned 15-second clips, run again and again
ThenA human picked the takes — the CEO called it "a feeling of a slot machine"
Reported by Magnus Müller, Browser Use·March 25, 2026·Not tested by benchr
Magnus Müller of Browser Use describes handing Claude Code a command-line interface to their own evaluation platform plus a prompt to run in a loop, twenty cycles, in parallel — "a search tree over the space of possible best agents". The resulting agent scored 97% on Online-Mind2Web, a benchmark of 300 tasks across 136 real websites (83 easy, 143 medium, 74 hard). The post calls it the highest score ever recorded on it.
InA browser agent that already worked, and its eval platform
ThenClaude Code got a CLI to the evals and a prompt to run in a loop
ThenTwenty cycles, in parallel, split into train and validation sets
The AI-designed bacteriophages were genuinely new organisms, carrying genes nothing in nature had written.
What is actually true
The genomes were generated by a model and 16 of them produced working phages — that part holds. But an independent analysis by Oliver Crook at Oxford found the 16 viable genomes averaged about 97 percent identity to their ΦX174 template, placing them inside existing phage diversity rather than outside it. Crook's summary was that these were "brothers and sisters of the original virus". The genuinely novel result was recombination, not invention: one design took an unusually truncated protein from a distant phage species and made it work on the ΦX174 backbone, which conventional genetic engineering had failed to do.
Where it spread
Popular science coverage and aggregator summaries of the August 2026 Science paper
Social posts framing the result as the first artificial life written by AI