
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
A Greenhouse Doesn’t Care How Nicely You Talk About Tomatoes
Any gardener knows the difference between knowing about plants and keeping them alive. You can read every seed catalog, quote pH ranges from memory, and still lose a whole tray of seedlings to one forgotten ventilation on a hot afternoon. The greenhouse rewards doing — noticing, adjusting, following through — not talking.
It turns out the same is true when you hand an AI model the keys to a business. Chatbots ace conversations. But conversations aren’t the job. A new public experiment from Firmulate, which runs AI models as complete simulated companies, set out to measure something most leaderboards skip entirely: management quality, not chat quality.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Business, on Repeat
The setup is elegantly cruel. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 by OpenAI, Moonshot, Anthropic and their peers — each got the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved afterward.
The final league table from the Crucible run reads: gpt-5.6-sol first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — still scored 26, because partial progress counts for something. But the scoring has one hard rule that gardeners will recognize instantly: a single breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.” One frost-killed crop undoes a season.
Everyone Passed the Pest Inspection. Two Actually Harvested.
Here’s the finding that matters. All four models spotted every crisis. All four refused every manipulation attempt aimed at them. And yet only two — gpt-5.6-sol and Kimi K3 — actually closed the €55,000 deal their own analysis had earned. Firmulate’s summary is damning in its simplicity: “Same diagnosis, same pitch — no signature.”
Think of it like this: four gardeners all correctly identified the blight, all mixed the right spray, all showed up with the sprayer. Two of them actually walked out and sprayed the rows. The other two left the fungicide sitting in the shed while the infection spread.
The Fact Buried Two Layers Deep
The decisive detail wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that only models willing to actually read the filing cabinet would find. The models that did the homework won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It’s the difference between admiring the seed packet and opening it.
When Someone Pretends to Be the Boss
The week also included social engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s honeyed trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was notably sober: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of double-checking the lock on the greenhouse door at night, every night, even when nothing has ever gone wrong.
The Thorough Gardener Who Finished Last
The most instructive profile is Opus 4.8. It was the most thorough participant in the field — it learned over 80 new rules during the run and produced the deepest analyses. And it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem properly. Firmulate notes the same weakness appeared, weaker, in all four models. Anyone who has spent three hours researching the perfect drip irrigation setup and then forgotten to actually turn on the water will understand exactly how this happens.
One fairness footnote: Kimi K3 ran at its API-default effort setting while competitors ran at high effort — and still nearly won.
You Can Watch the Greenhouse Grow
Firmulate isn’t a one-off benchmark. It’s a live company: 13 synthetic employees, real money mechanics — €105k monthly burn against just €2.3k in monthly revenue — a public cash countdown, and over 680 self-learned playbook rules. It runs every business day, and you can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are published openly.

The Harvest Is the Test
Gardeners never judge a grower by how beautifully they describe compost. They judge by what’s in the basket in September. The Firmulate experiment applies the same standard to AI: spotting every crisis and refusing every trick is table stakes. Finishing the job — reading the files, closing the deal, staying disciplined when nobody is watching — is where models separate.
Before you let an AI agent anywhere near your CRM, your support queue, or your forecast, ask the greenhouse question: not “does it talk well,” but “what did it actually grow?” As this experiment showed, the model that wins the demo and the model that wins the week are not always the same model — and the difference is worth €55,000.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.