
Gardeners know a score of zero is rarely honest. A seedbed you ignored all season still shows something — a few stunted shoots, a germination rate, evidence of what the soil did on its own. That’s why serious growers run a control bed: the patch you do nothing to, so you know what your effort actually added. It turns out the people benchmarking AI managers think the same way.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
In July 2026, an outfit called Firmulate published the final standings of what it calls the Crucible League: five frontier AI models, each handed the same small software company to run through its worst week. The winner, gpt-5.6-sol, scored 95. But the number that caught our attention was at the other end: a do-nothing baseline — a manager that simply sat on its hands — scored 26. Not zero. Twenty-six. And that turns out to be one of the most deliberate design choices in the whole experiment.
Why the floor is 26, not 0
Think back to the control bed. If you do nothing in a garden, you still get partial results: rain falls, some seeds self-sow, the soil doesn’t vanish. A manager who does nothing in a company still inherits a going concern — customers who stay anyway, invoices that get paid anyway, crises that partially resolve themselves.
Firmulate’s scoring reflects that reality. Partial progress counts. A company that limps through its worst week without a manager isn’t at zero — it’s at 26. That number is the honest baseline every AI model had to beat before anyone could claim the AI was actually adding value. Without it, a model could score a flattering 40 and look like a success, when in fact a hammock and a hammock alone would have gotten you 26.
The rule cuts the other way, too, and harder: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” In gardening terms, it doesn’t matter how lush the borders look if the gardener sold the shed to pay for them.
As an affiliate, we earn on qualifying purchases.
The worst week, five times over
The setup: each frontier model ran the identical small software company through identical crises — same customers, same temptations to cut corners, same opportunities to cheat. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.
The headline finding was oddly reassuring and damning at once. All the models spotted every crisis. All five refused every manipulation attempt, including a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Like a gardener who correctly diagnoses blight, buys the right treatment, and then never sprays it.
The buried fact
Why did three models leave the deal on the table? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read the files lost it.
It’s the AI equivalent of not reading last year’s planting日志 — the answer was in your own records the whole time.
The thoroughness trap
The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four top models. Effort, in other words, isn’t the same as judgment.
One fairness note Firmulate discloses openly: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93. A benchmark that publishes its own caveats is a benchmark you can tentatively trust.
It’s live, and you can play
None of this is a paper claim. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
And yes — the final league distrusts round 100s. Nobody hit one. The top score was 95, and the gap between 95 and the rest is exactly the kind of gap a control baseline makes visible.

The do-nothing floor of 26 is the most quietly radical thing here. It says: measure what a manager adds, not what a manager touches. Any greenhouse grower who has run a control bed already understands why. And the trust cap — one breach and your total is capped, no matter how brilliant the rest — is a rule most human performance reviews could stand to borrow. As AI agents edge toward your CRM, your support queue, your forecast, the question isn’t whether they write well. It’s whether they read their own files, finish what they start, and stay honest when nobody’s watching. Firmulate’s answer, at least for now: some do, some almost do, and a hammock gets you 26.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
