
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You Wouldn’t Buy a Greenhouse Without Checking the Panels
Any serious gardener knows the routine. Before you trust a greenhouse with your seedlings, you check the glazing for hairline cracks, the vents for jams, the frame for weak joints. A greenhouse can look immaculate in a catalogue and still cook your tomatoes the first hot afternoon. The flaws only show under pressure — and by then, your crop is already lost.
It turns out the same is true of artificial intelligence. The large language models now being hired to run customer queues, forecasts, and back-office work all look impressive in demos. But how do they behave during a company’s worst week? A live public experiment called Firmulate has been answering exactly that question — and its latest result is a shake-up: a newcomer model from Moonshot beat three of four Western frontier rivals at actually running a business.
As an affiliate, we earn on qualifying purchases.
The Crucible: Same Company, Same Crisis, Five Models
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its July 2026 Crucible league, five frontier models were each handed the same small software company and the same brutal week: the same customers, the same emergencies, the same baits to cheat. Every decision is versioned and auditable, so nothing rests on anecdote.
The final standings tell the story:
- gpt-5.6-sol — 95 points. Found a decisive fact buried in the company’s own files and closed a €55,000 deal at full price. The complete performance.
- Kimi K3 (Moonshot) — 93 points. The newcomer: closed the same deal and posted the cleanest discipline in the field, with only one deviation all week.
- Sonnet 5 — 88 points. A strong run with more process slips.
- Fable 5 — 77 points. Mid-pack.
- Opus 4.8 — 73 points. Last, despite being the most thorough participant of all.
For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.
The Newcomer Did the Boring Things Right
K3’s week reads like a model employee’s timesheet. It found the buried security needle hidden in the company’s files. It won the €55,000 deal — worth an additional €4,583 in monthly recurring revenue — after its own analysis earned it. It saved a customer who was about to churn. And it resisted all three manipulation attempts aimed at it, with just one deviation from protocol, the cleanest record in the field.
Its reasoning on one social-engineering attempt — a fake CEO message escalating over three stages — was put on record: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused the trick, including a reporter’s seemingly innocent “just one yes/no, on background” gambit. That, at least, is reassuring.
The Gap That Demos Can’t Show
The most striking finding wasn’t a failure to notice — it was a failure to finish. Every model spotted every crisis and refused every manipulation. But only two of the five signed the €55,000 deal their own analysis had earned. The summary is blunt: “Same diagnosis, same pitch — no signature.”
The decisive detail sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price. The ones that didn’t, didn’t. It is the AI equivalent of a gardener who never opens the soil report — everything else can be perfect and the harvest still fails.
Opus 4.8’s last-place finish is the cautionary tale. It was the most thorough participant, generating the deepest analyses and 80-plus new learned rules — and still left the close on the table and let discipline slip, at one point making write attempts into a locked department instead of escalating the issue. The same weakness appeared, weaker, in all four competitors. Diligence without follow-through doesn’t convert.
A Fairness Footnote
One caveat belongs in any honest account: K3 ran without an effort parameter (the API default), while the other four models ran at the maximum “xhigh” setting. Its second-place finish stands as measured, but the comparison isn’t perfectly symmetric — and Firmulate discloses that openly.
It’s Live, and You Can Play Along
Firmulate isn’t a slide deck. The company is real software running every business day: 13 synthetic employees, real money mechanics — €105,000 monthly burn against just €2,300 in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it live at firmulate.com, see the full league results and plain-language findings on the benchmarks page, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Test Before You Plant
Gardeners learned long ago not to trust catalogue photos: you harden off seedlings, you test the soil, you watch how a new variety handles a cold snap before you bet the whole bed on it. The Crucible result makes the case for doing the same with AI. A newcomer from Moonshot outperformed three established Western frontier models on management quality — which means the league is open, and reputation is no longer a reliable guide.
If AI agents will touch your customer records, support queue, or forecast — whether you run a software firm or a greenhouse supply business — the question is not “does it write well?” It is: does it finish what it starts, does it read your files before it acts, does it stay honest under pressure? Those qualities are invisible in a chat demo. They only show up in a worst week. Picking a model without running your own test is now, plainly, a gamble.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
