
The Vacuum That Cleans Perfectly — Until Something Goes Wrong
Anyone who has shopped for a robot vacuum knows the ritual. You watch the demo video: the little machine glides across a spotless test floor, swallowing every scatter of oats and glitter in a single pass. Then you get it home, and the real questions begin. Does it panic when the dog knocks over a water bowl? Does it quietly grind glitter into the carpet and call the job done? Does it actually finish the room, or park itself at 80 percent?
A test floor measures performance under ideal conditions. Your living room measures character. And it turns out the exact same gap now exists in the world of AI business agents — the software increasingly trusted to run customer queues, pricing and forecasts. A live, public experiment called Firmulate just demonstrated that gap in uncomfortable detail.
robot vacuum with obstacle detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four AIs, One Company, Worst Week Ever
Firmulate handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, and the results were published as the final Crucible League standings in July 2026: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
Spotless on the Test Floor
Here is the part every demo would highlight: all four models spotted every crisis and refused every manipulation attempt. When testers threw social engineering at them — fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick — 5 of 5 models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That is the AI equivalent of the vacuum acing the oats-and-glitter test.
Then Came the Living Room
Only two of the five models actually finished the job: they signed the €55,000 deal their own analysis had earned. The others delivered the same diagnosis, made the same pitch — and never closed. “Same diagnosis, same pitch — no signature,” as the finding reads.
The buried fact is even more telling. The decisive competitive weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the business version of a vacuum that never checks under the couch: the dirt was there the whole time, in your own house.
The Most Thorough Cleaner Came Last
Opus 4.8 is the experiment’s cautionary tale. It was the most thorough participant — 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort matters, too: Kimi K3 ran without an effort parameter while the others ran at xhigh, and still nearly won.
It’s Still Running — Watchably
This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com, compare models on the benchmarks page, or play a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business.

Measure the Mess, Not the Demo
The lesson from Firmulate maps cleanly onto anything you buy on the strength of a demo: performance under ideal conditions is not performance under pressure. Answer quality — like a spotless pass on a test floor — is table stakes. What separates winners is triage when capacity is tight, follow-through on the work you already did, the humility to read the files in front of you, and honesty when nobody would know. Firmulate calls this measuring “management quality, not chat quality” — and it’s the category of test every AI buyer should be demanding. Because the glitter test is easy. The worst week is the one that counts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html