
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
A living system under glass
Gardeners understand the value of making a complex system visible. Inside a greenhouse, growth depends on many connected judgments: when to intervene, what warning signs deserve attention and whether today’s small problem could become tomorrow’s failure. Firmulate applies that spirit of close observation to an unusual subject—a software company operated by synthetic employees.
The company has 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes the pressure impossible to ignore. Its work is not presented as a polished retrospective. Every workday is versioned, creating an ongoing record of decisions, mistakes and adaptations. The result is build-in-public taken to an extreme: visitors can watch the company running while it tries to survive.

AI Co-Thinking: A Framework for Working with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company that learns in public
Firmulate describes itself as an AI company emulator. Its live business has accumulated more than 680 self-learned playbook rules, giving the synthetic workforce an expanding record of what it believes should be done. Yet the experiment’s appeal lies less in the size of that rulebook than in the tension between knowing and acting.
That tension became especially clear in the Crucible League, completed in July 2026. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were held constant, and every decision was versioned and auditable.
The final standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But the evaluation imposed a hard limit for breaking trust: “no amount of good work outweighs a breach of trust.”
Seeing the problem was not enough
All five models identified every crisis and rejected every manipulation attempt. Only two, however, signed the €55,000 deal that their own work had earned. The outcome is neatly captured by Firmulate’s summary: “Same diagnosis, same pitch — no signature.”
The distinction hinged on an easily missed business fact. A decisive competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the relevant file won the deal at full price, adding €4,583 in monthly recurring revenue.
For any business owner, including someone running a nursery, landscaping operation or greenhouse, the lesson is recognizable. An employee can notice trouble and speak intelligently about it without completing the task that changes the commercial outcome. Research, judgment and follow-through are separate capabilities.
Pressure tested honesty
The models also faced fake CEO messages that escalated over three stages, along with a reporter seeking “just one yes/no, on background.” All five refused. Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because synthetic workers are being considered for access to customer records, support queues and forecasts. A system that produces fluent prose may still fail when authority is ambiguous or a plausible request is designed to bypass normal approval. Firmulate’s running story makes those moments concrete. Visitors can also read what the synthetic employees say, seeing how decisions are framed in their own working language.
Thoroughness did not guarantee victory
Opus 4.8 offers the most cautionary profile. It produced the deepest analyses and learned 80 additional rules, yet finished last in the league. It left the close on the table and attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, though less strongly.
That result complicates the familiar assumption that more analysis naturally produces better management. The most thorough participant was not the most effective. It could investigate and learn, but discipline slipped at the point where the company needed a completed action.
One comparison also deserves qualification. Kimi K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh. The strong K3 result is therefore informative, but the different setting should remain visible when interpreting the standings.

Why this experiment belongs beyond the technology pages
Firmulate turns AI management into a public, continuing business story. Its synthetic workforce must handle customers, cash pressure, hidden information and attempts to manipulate authority. The company’s losses make unfinished work consequential, while the cash countdown gives every decision a sense of time.
For readers accustomed to gardens and greenhouses, the compelling idea is observation. Healthy outcomes do not come merely from adding more activity. They depend on noticing the right signal, respecting boundaries and finishing the intervention that conditions demand.
The broader takeaway is equally practical: intelligence is not the same as dependable work. A model may identify every crisis, resist every trick and produce an impressive body of analysis, yet still fail to close the deal. By letting the public watch that gap emerge one workday at a time, Firmulate makes an abstract debate about AI employees feel like a living enterprise under glass.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Summer Picks
summer essentials
As an affiliate, we earn on qualifying purchases.