firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good cleaning depends on what happens after the dirt is spotted

Anyone who cares about floor care knows that recognizing a mess is not the same as removing it. A machine can detect debris, map a room and still leave dust along the baseboards. Firmulate has uncovered a strikingly similar gap in frontier artificial intelligence: models may understand a business crisis, explain the correct response and nevertheless fail to finish the job.

The public experiment placed each frontier model in charge of the same small software company during its worst week. Every participant faced the same customers, crises and temptations, with every decision versioned and auditable. The resulting management records now supply a quiz built from 242 real, unedited decisions. Readers see what a model actually did and try to identify its author.

That makes the exercise more revealing than a contest in polished writing. Its central question is whether management behavior has a recognizable signature: exhaustive or concise, decisive or hesitant, disciplined or prone to procedural slips.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed on the danger but diverged on execution

The broad competence was impressive. All models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”

The decisive information was not hiding in the customer event. It sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR. This is the business equivalent of cleaning beneath a movable rug: the visible surface does not contain the whole problem, and the most consequential detail may be somewhere an impatient operator never checks.

The final Crucible League results from July 2026 reflect those differences in follow-through:

  • gpt-5.6-sol finished first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

A do-nothing baseline scored 26 because partial progress counts. Firmulate also imposed a hard ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”

Pressure exposed discipline, not just intelligence

The company’s bad week included fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because the models were not simply being tested on whether they could produce plausible business language. They had to preserve trust while customers, authority figures and outsiders applied pressure. Across the field, resistance to manipulation was consistent; the meaningful separation appeared in research depth, operational discipline and completion.

The most thorough manager did not win

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is precisely why the quiz works as an interactive article. Readers are not guessing from artificial samples written to showcase a model’s voice. They are comparing real decisions made under identical conditions. One response may read like a dissertation; another may be terse; another may refuse unnecessary communication. Those stylistic clues become more meaningful when attached to concrete consequences.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. Its second-place finish should be read with that testing difference in view.

A company designed to make consequences visible

Firmulate’s live company contains 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to watch how management choices compound over time.

The setup turns abstract claims about AI agents into observable business behavior. It asks whether a model reads the company’s own material, protects confidential information, respects operational boundaries and completes commercially valuable work. For enterprises, Firmulate also offers the same wargame against a read-only export of their business; nothing writes back to real systems.

Infographic —
The findings at a glance — source: firmulate.com.

The practical lesson: inspect the finished floor

For cleaning enthusiasts, the analogy is straightforward. A vacuum should not be judged solely by its navigation map, motor specification or confident status message. The floor is the evidence. In the same way, an AI manager should not be judged only by the elegance of its diagnosis. The important questions are whether it found the buried fact, resisted pressure, followed procedure and completed the valuable action.

Firmulate’s experiment suggests that frontier models possess distinct, measurable management personalities. They may see the same crisis and recommend the same response, yet differ at the moment when analysis must become accountable action. The quiz makes those differences unusually accessible—and gives readers a chance to discover whether they can recognize a model by the managerial trail it leaves behind.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Best Air Purifier Routine for Homes Near Busy Roads

Purify your home air effectively near busy roads with a tailored routine that adapts to pollution levels—discover the secrets to healthier indoor air.

I Tested Every Type Of Portable Air Conditioner—These Are Worth Buying (And These Aren’t)

A comprehensive test of various portable air conditioners reveals which models deliver value, efficiency, and cooling performance for consumers.

Why Filter Life Drops Faster in Homes With Pets

The presence of pets accelerates filter clogging due to allergens and debris, and understanding how to slow this process can keep your system running smoothly.

Maintaining and Cleaning Air Purifier Filters Properly

Discover essential tips for maintaining and cleaning your air purifier filters properly to ensure optimal performance and air quality.