firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A robot vacuum can look impressive in a showroom and still miss the crumbs under the table. The real test comes when the room changes: a chair moves, a spill appears, or the usual route is blocked. AI agents face a similar challenge in business. A polished answer is one thing; handling a messy week of real decisions is another.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate puts that distinction on display. Its live experiment treats AI models as managers of the same small software company, with customers to keep, crises to handle and tempting shortcuts to resist. The results point to a question for any business considering AI: can a model do more than identify the right move? Can it follow through?

One company, one rough week

In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The contest’s trust rule was blunt: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models failed to notice trouble. All of them spotted every crisis and refused every manipulation attempt. The gap came at the finish: only two signed a €55,000 deal that their own analysis had earned. As the experiment put it, “Same diagnosis, same pitch — no signature.” Sound advice does not guarantee a completed task.

The clue was in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail makes the exercise feel less like a quiz with an obvious answer and more like a practical test of whether an AI can connect scattered information to a decision.

Firmulate also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The test therefore tracked both commercial judgment and whether models held a boundary when pressured.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness wrinkle in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also judge decisions for themselves through a “guess the model” quiz built from 242 real, unedited management decisions.

A company you can watch

The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k/month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The company is real as a live experiment, and visitors can watch it at firmulate.com.

For leaders, the appeal is the step beyond watching. Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The exercise can put company-specific crises and playbooks under pressure and produce a board report with model rankings and weak points. Nothing writes back to real systems. That boundary matters: a business can examine how an AI might act before granting it access to live tools.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to a business pilot

The cleaning lesson is familiar: performance in a controlled demo tells you less than performance across a whole home. Firmulate’s experiment suggests the same is true of AI at work. Models may spot the crisis and propose the pitch, yet still miss the buried clue or fail to close. A pilot using a read-only company export offers a way to see those strengths and gaps against your own business. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mold‑Fighting Purifiers: Uv‑C and Other Technologies

Mold-fighting purifiers use a combination of technologies like UV-C light, HEPA filters,…

What an AI’s Management Style Reveals When the Mess Gets Real

A live management wargame shows how frontier AI models differ in research, discipline and follow-through—even when they diagnose the same crisis.

What Is Cfm In Vacuums: What No One Told You (Vacuums & Floor Care Guide)

Just understanding CFM in vacuums reveals surprising insights into cleaning power that you won’t want to miss.

Dyson Gen5detect Review: Pros, Cons, and Who It’s For

Discover the Dyson Gen5detect vacuum’s features, pros, cons, and ideal users in this comprehensive review of top cordless vacuums.