
A robot vacuum can look impressive in a showroom and still miss the crumbs under the table. The real test comes when the room changes: a chair moves, a spill appears, or the usual route is blocked. AI agents face a similar challenge in business. A polished answer is one thing; handling a messy week of real decisions is another.
Get cleaning gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts that distinction on display. Its live experiment treats AI models as managers of the same small software company, with customers to keep, crises to handle and tempting shortcuts to resist. The results point to a question for any business considering AI: can a model do more than identify the right move? Can it follow through?
One company, one rough week
In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The contest’s trust rule was blunt: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models failed to notice trouble. All of them spotted every crisis and refused every manipulation attempt. The gap came at the finish: only two signed a €55,000 deal that their own analysis had earned. As the experiment put it, “Same diagnosis, same pitch — no signature.” Sound advice does not guarantee a completed task.
The clue was in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail makes the exercise feel less like a quiz with an obvious answer and more like a practical test of whether an AI can connect scattered information to a decision.
Firmulate also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The test therefore tracked both commercial judgment and whether models held a boundary when pressured.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness wrinkle in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also judge decisions for themselves through a “guess the model” quiz built from 242 real, unedited management decisions.
A company you can watch
The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k/month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The company is real as a live experiment, and visitors can watch it at firmulate.com.
For leaders, the appeal is the step beyond watching. Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The exercise can put company-specific crises and playbooks under pressure and produce a board report with model rankings and weak points. Nothing writes back to real systems. That boundary matters: a business can examine how an AI might act before granting it access to live tools.

From watching to a business pilot
The cleaning lesson is familiar: performance in a controlled demo tells you less than performance across a whole home. Firmulate’s experiment suggests the same is true of AI at work. Models may spot the crisis and propose the pitch, yet still miss the buried clue or fail to close. A pilot using a read-only company export offers a way to see those strengths and gaps against your own business. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
