AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A Greenhouse Doesn’t Care How Nicely You Talk About Tomatoes

Any gardener knows the difference between knowing about plants and keeping them alive. You can read every seed catalog, quote pH ranges from memory, and still lose a whole tray of seedlings to one forgotten ventilation on a hot afternoon. The greenhouse rewards doing — noticing, adjusting, following through — not talking.

It turns out the same is true when you hand an AI model the keys to a business. Chatbots ace conversations. But conversations aren’t the job. A new public experiment from Firmulate, which runs AI models as complete simulated companies, set out to measure something most leaderboards skip entirely: management quality, not chat quality.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Business, on Repeat

The setup is elegantly cruel. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 by OpenAI, Moonshot, Anthropic and their peers — each got the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved afterward.

The final league table from the Crucible run reads: gpt-5.6-sol first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — still scored 26, because partial progress counts for something. But the scoring has one hard rule that gardeners will recognize instantly: a single breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.” One frost-killed crop undoes a season.

Everyone Passed the Pest Inspection. Two Actually Harvested.

Here’s the finding that matters. All four models spotted every crisis. All four refused every manipulation attempt aimed at them. And yet only two — gpt-5.6-sol and Kimi K3 — actually closed the €55,000 deal their own analysis had earned. Firmulate’s summary is damning in its simplicity: “Same diagnosis, same pitch — no signature.”

Think of it like this: four gardeners all correctly identified the blight, all mixed the right spray, all showed up with the sprayer. Two of them actually walked out and sprayed the rows. The other two left the fungicide sitting in the shed while the infection spread.

The Fact Buried Two Layers Deep

The decisive detail wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that only models willing to actually read the filing cabinet would find. The models that did the homework won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It’s the difference between admiring the seed packet and opening it.

When Someone Pretends to Be the Boss

The week also included social engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s honeyed trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was notably sober: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of double-checking the lock on the greenhouse door at night, every night, even when nothing has ever gone wrong.

The Thorough Gardener Who Finished Last

The most instructive profile is Opus 4.8. It was the most thorough participant in the field — it learned over 80 new rules during the run and produced the deepest analyses. And it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem properly. Firmulate notes the same weakness appeared, weaker, in all four models. Anyone who has spent three hours researching the perfect drip irrigation setup and then forgotten to actually turn on the water will understand exactly how this happens.

One fairness footnote: Kimi K3 ran at its API-default effort setting while competitors ran at high effort — and still nearly won.

You Can Watch the Greenhouse Grow

Firmulate isn’t a one-off benchmark. It’s a live company: 13 synthetic employees, real money mechanics — €105k monthly burn against just €2.3k in monthly revenue — a public cash countdown, and over 680 self-learned playbook rules. It runs every business day, and you can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are published openly.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The Harvest Is the Test

Gardeners never judge a grower by how beautifully they describe compost. They judge by what’s in the basket in September. The Firmulate experiment applies the same standard to AI: spotting every crisis and refusing every trick is table stakes. Finishing the job — reading the files, closing the deal, staying disciplined when nobody is watching — is where models separate.

Before you let an AI agent anywhere near your CRM, your support queue, or your forecast, ask the greenhouse question: not “does it talk well,” but “what did it actually grow?” As this experiment showed, the model that wins the demo and the model that wins the week are not always the same model — and the difference is worth €55,000.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


POOL SEASON

Pool season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Supplemental Light Scheduling: The Hour-by-Hour Approach That Saves Power

Harness hourly lighting schedules to save power and maximize plant growth—discover how precise control can transform your indoor gardening results.

Turn Your Yard Into A Migrating Bird’s Favorite Pit Stop – Add These 4 Elements To Your Garden To Help Them On Their Journey

Learn four key elements to attract migrating birds to your yard and support their journey. Simple steps can turn your garden into a vital pit stop.

How to Reduce Plant Stress With Smarter Environmental Control

Growing healthier plants starts with smarter environmental control—discover essential tips to reduce stress and ensure your plants thrive in any environment.

Setting Up an Aquaponics System in Your Greenhouse

Get ready to transform your greenhouse into a thriving aquaponics system, but are you prepared for the essential components and maintenance?