AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Could an AI keep a greenhouse business steady through its worst week?

A greenhouse operator might face a customer threatening to leave, a rival undercutting a key product and a convincing message that appears to come from the CEO. Those pressures rarely arrive one at a time. Firmulate’s live experiment puts AI models in charge of a small software company facing that kind of pileup, offering garden and outdoor-living businesses a way to think about what agentic AI might do when decisions carry real consequences.

Same company, same crises

In the final Crucible League, published in July 2026, frontier models ran the same small company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The point was to observe management under pressure, not to judge how polished a model’s chat sounded.

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” For a greenhouse business, the equivalent could be an AI agent that correctly identifies a promising wholesale customer or supply risk, then fails to carry the decision through.

The standings show a wide spread. gpt-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.” K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

The clue was already in the files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding points to a practical challenge for any business considering AI: useful information may already exist in its records, but success depends on whether an agent finds and acts on it.

The test also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee follow-through

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. The lesson is concrete: capable analysis does not automatically translate into reliable execution or respect for business boundaries.

Firmulate’s live company makes the exercise watchable. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.

From watching to your own pilot

For a grower, greenhouse supplier or outdoor-living business, the experiment raises a timely question: how would an AI agent handle your customer records, seasonal pressures, pricing decisions and internal rules? Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. The exercise tests crisis scenarios and playbooks without writing back to real systems.

See the live experiment at Firmulate, or try the guess-the-model quiz.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

The league suggests that recognizing a crisis is only part of the job. Finding buried evidence, closing a deal and staying within authority matter too. Enterprises can explore those questions with a pilot using a read-only export of their own business; nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Summer Entertaining with the Ninja CREAMi Deluxe Ice Cream Maker & Accessories

Discover must-have accessories for your Ninja CREAMi Deluxe to elevate summer gatherings with delicious frozen treats and creative pairings.

Dutch Bucket Basics: How to Set Up Drainage So Plants Don’t Drown

Keen to master Dutch bucket drainage? Discover essential tips to prevent overwatering and ensure healthy plant growth.

Summertime Feast with the Ninja Foodi 10 Quart DualZone Air Fryer

Discover easy summer recipes using the Ninja Foodi XL 2-Basket Air Fryer—perfect for delicious family meals and outdoor entertaining.

Before You Let AI Run the Greenhouse, Watch It Fail a Stress Test First

A newcomer AI from Moonshot beat three of four Western frontier models at running a company under pressure — and the league table is suddenly wide open.