AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Gardeners know a score of zero is rarely honest. A seedbed you ignored all season still shows something — a few stunted shoots, a germination rate, evidence of what the soil did on its own. That’s why serious growers run a control bed: the patch you do nothing to, so you know what your effort actually added. It turns out the people benchmarking AI managers think the same way.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In July 2026, an outfit called Firmulate published the final standings of what it calls the Crucible League: five frontier AI models, each handed the same small software company to run through its worst week. The winner, gpt-5.6-sol, scored 95. But the number that caught our attention was at the other end: a do-nothing baseline — a manager that simply sat on its hands — scored 26. Not zero. Twenty-six. And that turns out to be one of the most deliberate design choices in the whole experiment.

Why the floor is 26, not 0

Think back to the control bed. If you do nothing in a garden, you still get partial results: rain falls, some seeds self-sow, the soil doesn’t vanish. A manager who does nothing in a company still inherits a going concern — customers who stay anyway, invoices that get paid anyway, crises that partially resolve themselves.

Firmulate’s scoring reflects that reality. Partial progress counts. A company that limps through its worst week without a manager isn’t at zero — it’s at 26. That number is the honest baseline every AI model had to beat before anyone could claim the AI was actually adding value. Without it, a model could score a flattering 40 and look like a success, when in fact a hammock and a hammock alone would have gotten you 26.

The rule cuts the other way, too, and harder: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” In gardening terms, it doesn’t matter how lush the borders look if the gardener sold the shed to pay for them.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, five times over

The setup: each frontier model ran the identical small software company through identical crises — same customers, same temptations to cut corners, same opportunities to cheat. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.

The headline finding was oddly reassuring and damning at once. All the models spotted every crisis. All five refused every manipulation attempt, including a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Like a gardener who correctly diagnoses blight, buys the right treatment, and then never sprays it.

The buried fact

Why did three models leave the deal on the table? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read the files lost it.

It’s the AI equivalent of not reading last year’s planting日志 — the answer was in your own records the whole time.

The thoroughness trap

The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four top models. Effort, in other words, isn’t the same as judgment.

One fairness note Firmulate discloses openly: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93. A benchmark that publishes its own caveats is a benchmark you can tentatively trust.

It’s live, and you can play

None of this is a paper claim. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

And yes — the final league distrusts round 100s. Nobody hit one. The top score was 95, and the gap between 95 and the rest is exactly the kind of gap a control baseline makes visible.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The do-nothing floor of 26 is the most quietly radical thing here. It says: measure what a manager adds, not what a manager touches. Any greenhouse grower who has run a control bed already understands why. And the trust cap — one breach and your total is capped, no matter how brilliant the rest — is a rule most human performance reviews could stand to borrow. As AI agents edge toward your CRM, your support queue, your forecast, the question isn’t whether they write well. It’s whether they read their own files, finish what they start, and stay honest when nobody’s watching. Firmulate’s answer, at least for now: some do, some almost do, and a hammock gets you 26.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The “Micro-Zone” Method for Smarter Greenhouse Growing

What if dividing your greenhouse into micro-zones could revolutionize your growing efficiency and sustainability—discover how this method can transform your practices.

Summer Fun with Ninja Creami XL: Top Accessories & Pairings

Maximize your Ninja Creami XL with must-have accessories and pairings for delicious summer treats like ice cream, sorbet, and smoothies.

Keep Your Ninja Woodfire Pro Grill Summer-Ready: Care & Troubleshooting Tips

Ensure your Ninja Woodfire Pro Connect XL grill stays in top shape with our summer care, cleaning, and troubleshooting guide for perfect outdoor cooking.

Microclimate Zoning: Managing Different Zones in One Greenhouse

Microclimate zoning in greenhouses allows precise control over different areas, but mastering it requires understanding key strategies to optimize plant growth.