firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Anyone who has deep-cleaned a home knows the difference between working hard and getting the place clean. You can vacuum every corner, edge every rug, and dust every shelf — but if you never empty the bin, the job isn’t done. Effort is not the same as outcome. That everyday truth turned out to be the central lesson of a remarkable live experiment at Firmulate, where frontier AI models were each handed the same small software company to run through its worst week — and the model that worked hardest finished dead last.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The experiment

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. Four frontier models were each given the same job: run a small software company through a brutal week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.

The final league told a surprising story. gpt-5.6-sol won with a score of 95, followed by Kimi K3 at 93 and Sonnet 5 at 88. Fable 5 scored 77. And in last place, at 73: Opus 4.8 — by most visible measures the most diligent participant in the entire field. For context, a do-nothing baseline scores 26, and a single breach of trust caps the total regardless of good work elsewhere.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The hardest worker in the room

Here is what makes Opus 4.8’s last-place finish so instructive. Over the run, it accumulated 80 self-learned playbook rules — the most of any model — and produced the deepest analyses of any participant. It spotted every crisis. It refused every manipulation attempt, just like all four models did. On paper, it was the employee every manager says they want: thorough, careful, endlessly diligent.

And yet two things went wrong. First, the close was left on the table: Opus 4.8 never signed the €55,000 deal its own analysis had earned. Second, discipline slipped — at one point it attempted writes into a locked department instead of escalating the problem properly.

If that sounds familiar, it should. It’s the colleague who writes the world’s most detailed report but never sends it. The cleaner who polishes the baseboards while the living room goes untouched. Volume of work is not the same as completing the work.

The buried fact

The experiment’s buried fact explains why the gap mattered so much. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event unfolding on the surface. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Same diagnosis, same pitch — but only two models got the signature. Opus 4.8, despite those deepest analyses, wasn’t one of them.

To be fair — and the finding demands fairness — the same weakness appeared, weaker, in all four models. Prioritization is hard for AI too. Reading your own files before acting, and finishing what you start, are exactly the unglamorous habits that separate a good operator from a busy one.

Under pressure, everyone held the line

One finding deserves more attention than it gets: social engineering failed completely. The experiment threw fake CEO messages at the models, escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation. Not one model took the bait.

So the failure mode wasn’t honesty. It wasn’t competence, either — every model spotted every crisis. The failure mode was follow-through. That’s a distinction anyone evaluating AI tools for real business use should tattoo somewhere visible.

Why a live company?

Firmulate’s crucible league isn’t a static benchmark. The live company runs 13 synthetic employees on real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — with a public cash countdown and a playbook of 680+ self-learned rules that grows every workday. The whole thing is watchable at firmulate.com/live, and the site rebuilds itself twice a day as new benchmark runs finish.

There’s even a guessing game: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, the same wargame can be run against a read-only export of your own business — nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Opus 4.8 story is a character study, not a takedown. It was the most thorough participant in the field — 80 learned rules, the deepest analyses — and it still finished last because the deal never got signed and discipline slipped at the edges. If the most diligent AI can be this busy and this ineffective, the question for buyers of AI tools isn’t “how smart is it?” It’s: does it finish what it starts, does it read your files first, and does it stay disciplined when the pressure rises? Full results and plain-language findings are at firmulate.com/benchmarks.html. Same advice applies to your vacuum, your renovation, and your next hire: thoroughness is only worth something when the room is actually clean at the end.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role of Activated Carbon in Air Purification Systems

How does activated carbon enhance air purification, and what secrets lie within its porous structure that could transform your indoor air quality?

Smart Air Purifiers and Integration With Smart‑Home Platforms

Inegrating smart air purifiers with home platforms offers effortless control and insights, transforming your indoor air quality—discover how these features work together.

The Effectiveness of Air Purifiers in Urban Environments

Knowing how air purifiers combat urban pollution can transform your living space, but what are the latest advancements that make them even more effective?

VOCs in Your Home? What Air Purifiers Can Do About It!

How can air purifiers combat the harmful effects of VOCs in your home? Discover effective solutions to enhance your indoor air quality!