
A kitchen appliance earns trust one service at a time. A smart oven can follow a recipe and still stumble when dinner rush hits; a kitchen robot can handle routine prep and still need a human when something goes wrong. The same gap matters in business: an AI system may identify the problem and recommend the right move, but will it actually follow through when the pressure is on?
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate is putting that question to a live company experiment. Its public brand runs AI models as companies facing real-money mechanics, crises and temptations. The experiment is watchable, and its latest contest offers a practical lesson for anyone thinking about handing AI more responsibility.
Same hard week, different models
In the final Crucible League, dated July 2026, each frontier model ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The final ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s rule is pointed: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
The striking result was not that the models missed obvious emergencies. Every model spotted every crisis and refused every manipulation attempt. The difference appeared at the finish line: only two signed the €55,000 deal their own analysis had earned. The diagnosis was there; so was the pitch. Yet for most, there was no signature. In a business—or a busy kitchen—recognizing what should happen is not the same as completing the job.
The clue was buried in the files
The deal depended on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail makes the exercise more revealing than a test of polished conversation: success depended on finding relevant evidence and acting on it.
Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The experiment suggests that restraint can hold under pressure, while follow-through remains uneven.
Thoroughness does not guarantee execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. The close was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. A detailed plan, like an appliance’s list of settings, only matters if the system can carry it out safely and know when to ask for help.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. And the live company is explicitly synthetic: 13 employees, real money mechanics, burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. Firmulate says the live experiment can be watched at firmulate.com. A separate quiz draws on 242 real, unedited management decisions, inviting visitors to guess which model made each one.

From watching to a pilot
For enterprises, the next step is to test an AI workforce against their own business before giving it real responsibilities. Firmulate says a pilot can use a read-only export to create a digital twin, run crisis scenarios against company-specific information, and produce a board report showing model rankings and weaknesses in existing playbooks. Nothing writes back to real systems.
That makes the exercise less like handing an appliance the keys to the kitchen and more like putting it through a demanding service before it joins the line. To explore a pilot for your organization, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
