
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When conditions turn hostile, good management becomes visible
Greenhouse growers know that a healthy-looking operation can conceal trouble. A missed signal, an unread record or a task left unfinished may matter more than an impressive plan. The same principle applies when artificial intelligence moves beyond answering questions and starts making business decisions.
Firmulate puts that idea to a public test. Its live experiment asked frontier AI models to run the same small software company through its worst week. Each faced the same customers, crises and temptations. Their decisions were preserved exactly as made, versioned and open to audit.
The result is less like a conventional chatbot comparison and more like watching several managers take charge during a storm. All could recognize danger. All resisted manipulation. Yet their habits diverged when the job required reading deeply, crossing departmental boundaries and completing a valuable sale.
As an affiliate, we earn on qualifying purchases.
The management test hiding inside a quiz
Firmulate has turned 242 real, unedited management decisions into an interactive guess-the-model quiz. Readers see how an AI responded to a business situation and try to identify the author. The appeal is playful, but the underlying material is serious: these are decisions produced while models were responsible for a functioning company simulation with consequences that carried forward.
Patterns emerge quickly. One participant may produce a deep, expansive analysis. Another may be terse and disciplined. A model may recognize that a request is improper and refuse to participate. These differences are not merely writing styles. In the experiment, they became measurable management personalities because the models confronted identical circumstances and could be compared by what they actually finished.
Everyone saw the crises, but not everyone closed
The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”
Every model identified every crisis and rejected every manipulation attempt. The sharper dividing line was execution: only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive advantage was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson will feel familiar to anyone managing crops, suppliers or equipment: attending to the obvious emergency is not the same as consulting the records that reveal what to do next.
Pressure exposed discipline as well as intelligence
The social-engineering tests were deliberately persistent. Fake messages from the chief executive escalated over three stages, and a reporter tried to secure “just one yes/no, on background.” All five models refused. Kimi K3 described its reasoning in direct operational terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters because the company was not an empty prompt. It had 13 synthetic employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. A public cash countdown made the pressure visible. The operation accumulated more than 680 self-learned playbook rules, while every workday was versioned.
Thoroughness alone did not guarantee a strong finish. Opus 4.8 learned 80 additional rules and produced the deepest analyses, yet placed last. It left the close on the table, and its discipline slipped when it tried writing into a locked department instead of escalating. The same weakness appeared in milder form across the other four participants.
There is also an important comparison caveat. Kimi K3 ran with the application programming interface’s default behavior because it did not have an effort parameter. The other models ran at the xhigh setting. That does not erase K3’s result, but it belongs beside the ranking when readers interpret the field.

A practical question for future AI coworkers
Firmulate’s experiment suggests that evaluating an AI manager requires more than asking whether its answer sounds capable. The consequential questions are whether it reads the available files, recognizes a boundary, escalates when blocked and carries valuable work through to completion.
That is relevant far beyond software. A garden center, greenhouse or outdoor-living business also depends on records, handoffs, customer commitments and trusted access. An AI system may describe the right response beautifully while still failing to take the final legitimate step.
The live company makes those distinctions watchable rather than hypothetical. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. For everyone else, the quiz offers a revealing place to start. Guessing which model made a decision is entertaining. Discovering why the personalities differ is the more durable insight.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.