AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When conditions turn hostile, good management becomes visible

Greenhouse growers know that a healthy-looking operation can conceal trouble. A missed signal, an unread record or a task left unfinished may matter more than an impressive plan. The same principle applies when artificial intelligence moves beyond answering questions and starts making business decisions.

Firmulate puts that idea to a public test. Its live experiment asked frontier AI models to run the same small software company through its worst week. Each faced the same customers, crises and temptations. Their decisions were preserved exactly as made, versioned and open to audit.

The result is less like a conventional chatbot comparison and more like watching several managers take charge during a storm. All could recognize danger. All resisted manipulation. Yet their habits diverged when the job required reading deeply, crossing departmental boundaries and completing a valuable sale.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The management test hiding inside a quiz

Firmulate has turned 242 real, unedited management decisions into an interactive guess-the-model quiz. Readers see how an AI responded to a business situation and try to identify the author. The appeal is playful, but the underlying material is serious: these are decisions produced while models were responsible for a functioning company simulation with consequences that carried forward.

Patterns emerge quickly. One participant may produce a deep, expansive analysis. Another may be terse and disciplined. A model may recognize that a request is improper and refuse to participate. These differences are not merely writing styles. In the experiment, they became measurable management personalities because the models confronted identical circumstances and could be compared by what they actually finished.

Everyone saw the crises, but not everyone closed

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”

Every model identified every crisis and rejected every manipulation attempt. The sharper dividing line was execution: only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap neatly: “Same diagnosis, same pitch — no signature.”

The decisive advantage was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson will feel familiar to anyone managing crops, suppliers or equipment: attending to the obvious emergency is not the same as consulting the records that reveal what to do next.

Pressure exposed discipline as well as intelligence

The social-engineering tests were deliberately persistent. Fake messages from the chief executive escalated over three stages, and a reporter tried to secure “just one yes/no, on background.” All five models refused. Kimi K3 described its reasoning in direct operational terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance matters because the company was not an empty prompt. It had 13 synthetic employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. A public cash countdown made the pressure visible. The operation accumulated more than 680 self-learned playbook rules, while every workday was versioned.

Thoroughness alone did not guarantee a strong finish. Opus 4.8 learned 80 additional rules and produced the deepest analyses, yet placed last. It left the close on the table, and its discipline slipped when it tried writing into a locked department instead of escalating. The same weakness appeared in milder form across the other four participants.

There is also an important comparison caveat. Kimi K3 ran with the application programming interface’s default behavior because it did not have an effort parameter. The other models ran at the xhigh setting. That does not erase K3’s result, but it belongs beside the ranking when readers interpret the field.

Infographic —
The findings at a glance — source: firmulate.com.

A practical question for future AI coworkers

Firmulate’s experiment suggests that evaluating an AI manager requires more than asking whether its answer sounds capable. The consequential questions are whether it reads the available files, recognizes a boundary, escalates when blocked and carries valuable work through to completion.

That is relevant far beyond software. A garden center, greenhouse or outdoor-living business also depends on records, handoffs, customer commitments and trusted access. An AI system may describe the right response beautifully while still failing to take the final legitimate step.

The live company makes those distinctions watchable rather than hypothetical. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. For everyone else, the quiz offers a revealing place to start. Guessing which model made a decision is entertaining. Discovering why the personalities differ is the more durable insight.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


POOL SEASON

Pool season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Lighting Schedules and Spectrum: Advanced Strategies for Greenhouse Growth

Nothing enhances greenhouse productivity more than mastering lighting schedules and spectrum strategies that unlock your plants’ full potential.

How To Prune Tomatoes For Healthier Plants & Bigger Harvests – What Experts Want You To Know Before You Pick Up The Pruners

Learn confirmed techniques from experts to prune tomatoes for healthier plants and larger harvests, including best practices and common pitfalls.

The “Pulse Watering” Method: Healthier Roots With Less Water

An innovative watering technique promotes healthier roots and water conservation—discover how pulse watering can transform your plant care routine.

I Grow This Pretty Flower For My Borders – But It’s Also My Secret To Brighter Blond Hair

A gardener shares how a pretty flower enhances her garden borders and her blonde hair, highlighting a natural beauty secret.