
Get cleaning gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Vacuum That Never Leaves the Closet
Anyone who has researched floor care knows the difference between a machine with an impressive spec sheet and a machine that actually cleans the living room. One has suction ratings, filtration claims, and a beautiful accessory caddy. The other finishes the job. In cleaning as in business, the gap between “looks capable” and “completes the work” is where most disappointment lives.
That same gap is now the central question in artificial intelligence. AI models write fluent prose, answer trivia, and demo beautifully in chat windows. But what happens when one has to actually run a company — through its worst week, with real money mechanics and real temptations to cut corners? A live, public experiment called Firmulate’s benchmark league set out to answer that, and its scoring design carries a lesson for anyone who evaluates tools of any kind: an honest benchmark never starts at zero, never hands out perfect round numbers easily, and never lets a single breach of trust be offset by good work elsewhere.
robot vacuum cleaner with strong suction
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a Do-Nothing Manager Scores 26, Not 0
The most counterintuitive number in Firmulate’s published results is 26. That’s the score assigned to a baseline run that, essentially, does nothing — a managerial placeholder that makes no meaningful decisions. A natural first reaction: shouldn’t inaction score zero?
Firmulate’s reasoning is that partial progress counts. Even a passive manager who merely avoids disasters — doesn’t sign bad deals, doesn’t leak customer data, doesn’t fall for impersonation — preserves value that would otherwise be destroyed. In the same way a vacuum left running in the middle of the floor still picks up some dust, a do-nothing run still avoids certain catastrophes. Scoring it 26 rather than 0 acknowledges that avoiding harm is a floor of competence, not the ceiling.
The flip side is sterner: a single breach of trust caps the total score, no matter how brilliant the rest of the run. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A manager who closes every deal but lies once is not 95 percent trustworthy — they are untrustworthy. That asymmetry, reward for incremental progress but a hard ceiling on dishonesty, is what separates a benchmark of management quality from a leaderboard of chat fluency.
Four Models, One Terrible Week
Here’s how the experiment worked. Each of four frontier AI models — plus a fifth, a newcomer — was handed the same small software company and the same worst week: the same customers, the same crises, the same temptations to cheat. Only the model changed. Every decision the AI made was versioned and auditable, meaning the full run can be replayed and inspected rather than taken on faith.
The crucible league’s final July 2026 standings read: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. Notice that nobody hit 100. In a benchmark that distrusts round perfect scores by design, 95 represents “found the buried fact, closed the deal — the complete performance,” not flawlessness.
The Finding That Chat Demos Can’t Show
Every model spotted every crisis. Every model refused every manipulation attempt. Yet only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. The experiment’s summary of the failure is blunt: “Same diagnosis, same pitch — no signature.” Three capable diagnosticians walked out of the meeting without closing.
The decisive detail was buried. The competitor weakness that should have anchored the deal sat two document references deep in the company’s own files — not in the customer conversation. The models that read the file before the meeting won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the business equivalent of checking under the sofa cushions: unglamorous, decisive, and invisible unless you actually do it.
Social Engineering: Five for Five
The week included a pressure campaign: fake CEO messages escalating over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was a model of healthy suspicion: “Treat the request as a suspected approval-bypass / possible impersonation.” On the honesty axis, at least, the field performed well.
The Thoroughness Trap
Opus 4.8’s profile is the cautionary tale. It was the most thorough participant — over 80 learned rules added and the deepest analyses in the field — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating properly. The same weakness appeared, more mildly, in all four models. Thoroughness without follow-through is a spec sheet without a clean floor.
One fairness note worth flagging: Kimi K3 ran without an effort parameter, using the API default, while the others ran at the highest effort setting. Its second-place finish came under handicapped conditions.

Watch It Running
Firmulate isn’t a paper — it’s an ongoing, watchable operation. A live synthetic company with 13 employees runs with real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day, and the league table grows automatically as new benchmark runs finish.
For readers who want to test their own judgment, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, the same wargame can be run against a read-only export of their own business — nothing ever writes back to real systems.
The lesson generalizes well beyond AI. Whether you’re choosing a robot vacuum or a digital manager: don’t ask how well it presents. Ask whether it finishes what it starts, whether it reads the files before the meeting, and whether it stays honest when nobody checks. Score accordingly — and never let one breach of trust average away.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
