
A kitchen timer is good at one thing: telling you when time is up. Choosing an AI to run parts of a business is harder. A polished answer may sound promising, but will the system read the files, protect customers and finish the job? Firmulate put several frontier models through the same rough week at a software company to find out.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The test was about decisions, not chat
In Firmulate’s live experiment, each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workday is versioned, and its playbook has accumulated more than 680 self-learned rules.
The final July 2026 league table put gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The difference was in the follow-through
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive clue was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file found the competitor’s weakness and won the deal at full price, worth €4,583 in monthly recurring revenue.
That gap between identifying the right move and completing it is the story behind K3’s second-place result. The newcomer beat three of the four Western models in the field. K3 found the buried security needle, won the deal, saved the churning customer and resisted all three baits, with just one deviation. Firmulate describes its discipline as the cleanest in the field. The top spot was close: gpt-5.6-sol finished two points ahead.
The pressure included fake CEO messages escalating through three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
More effort did not guarantee a better finish
Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses, yet it placed last. It left the deal unsigned and slipped on discipline by attempting writes into a locked department instead of escalating. Firmulate says the same weakness appeared, more mildly, across all four Western models. The result is a useful reminder that extensive analysis and a strong-sounding diagnosis do not guarantee a completed, well-handled decision.
There is a fairness detail for readers weighing the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
A test you can watch
Firmulate presents this as a real, watchable experiment, not a slide deck. Readers can follow the company at Firmulate and review plain-language benchmark results at the benchmark page. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
For businesses considering AI agents in a CRM, support queue or forecast, the comparison raises a practical question: how will a model behave in your own working conditions? Firmulate says enterprises can run the wargame against a read-only export of their business, with no writes back to real systems.

Test the finish, not just the answer
Kimi K3’s near-top score shows the league is open, while the unsigned deals show why a general ranking cannot settle every buying decision. A model must do more than spot the right action: it must follow through and keep trust intact. Picking one without testing it against your own work is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
