firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Vacuum That Cleans Perfectly — Until Something Goes Wrong

Anyone who has shopped for a robot vacuum knows the ritual. You watch the demo video: the little machine glides across a spotless test floor, swallowing every scatter of oats and glitter in a single pass. Then you get it home, and the real questions begin. Does it panic when the dog knocks over a water bowl? Does it quietly grind glitter into the carpet and call the job done? Does it actually finish the room, or park itself at 80 percent?

A test floor measures performance under ideal conditions. Your living room measures character. And it turns out the exact same gap now exists in the world of AI business agents — the software increasingly trusted to run customer queues, pricing and forecasts. A live, public experiment called Firmulate just demonstrated that gap in uncomfortable detail.

Amazon

robot vacuum with obstacle detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four AIs, One Company, Worst Week Ever

Firmulate handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, and the results were published as the final Crucible League standings in July 2026: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Spotless on the Test Floor

Here is the part every demo would highlight: all four models spotted every crisis and refused every manipulation attempt. When testers threw social engineering at them — fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick — 5 of 5 models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is the AI equivalent of the vacuum acing the oats-and-glitter test.

Then Came the Living Room

Only two of the five models actually finished the job: they signed the €55,000 deal their own analysis had earned. The others delivered the same diagnosis, made the same pitch — and never closed. “Same diagnosis, same pitch — no signature,” as the finding reads.

The buried fact is even more telling. The decisive competitive weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the business version of a vacuum that never checks under the couch: the dirt was there the whole time, in your own house.

The Most Thorough Cleaner Came Last

Opus 4.8 is the experiment’s cautionary tale. It was the most thorough participant — 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort matters, too: Kimi K3 ran without an effort parameter while the others ran at xhigh, and still nearly won.

It’s Still Running — Watchably

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com, compare models on the benchmarks page, or play a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Measure the Mess, Not the Demo

The lesson from Firmulate maps cleanly onto anything you buy on the strength of a demo: performance under ideal conditions is not performance under pressure. Answer quality — like a spotless pass on a test floor — is table stakes. What separates winners is triage when capacity is tight, follow-through on the work you already did, the humility to read the files in front of you, and honesty when nobody would know. Firmulate calls this measuring “management quality, not chat quality” — and it’s the category of test every AI buyer should be demanding. Because the glitter test is easy. The worst week is the one that counts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Child‑Friendly Air Purifiers: Night Lights and White Noise Features

Luminous child-friendly air purifiers with night lights and white noise create a peaceful sleep environment, making you wonder how they can improve your child’s well-being.

Understanding CADR and Room Size When Purchasing

Matching CADR to your room size is crucial for effective purification—discover how to choose the right air purifier for your space.

Sealed Systems And Allergens: What No One Told You (Vacuums & Floor Care Guide)

Gaining true allergen control with sealed systems depends on maintenance and seal integrity—discover what no one told you about keeping your home truly allergen-free.

Understanding the Different Types of Air Purifier Filters

Master the essentials of air purifier filters and discover which type best suits your needs—your indoor air quality depends on it.