
Imagine hiring an assistant who can read your most secret files — not just to answer questions, but to truly understand your business. In a recent live experiment, AI models were put through a simulated crisis week, and the results could reshape how companies choose their AI helpers. The key was not just what they said, but what they read beneath the surface.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Virtual Business War
Four leading AI models, including GPT-5.6 and Kimi K3, faced the same challenge: manage a small software company’s worst week. This meant navigating customer crises, avoiding manipulation, and making critical decisions — all within a controlled, auditable environment. The goal? See if these AI models could spot hidden risks and act with integrity.
What the Models Saw and Did
Every model successfully identified each crisis and refused manipulative tactics like fake CEO messages or reporter tricks. Yet, only two of them actually closed the deal, signing a €55,000 contract based on their analysis. The rest, despite recognizing the problems, left the deal on the table or slipped into process slips.
As an affiliate, we earn on qualifying purchases.
The Hidden Key: Reading Beneath the Surface
The crucial difference? The winning models read deeper into the company’s own files — references two layers beneath the surface. This buried fact, hidden in internal documents, was decisive. The models that uncovered it won the contract, worth an additional €4,583 monthly recurring revenue.
Why This Matters for Business AI
If your AI assistant only responds based on surface-level data or chat prompts, it might miss the critical details that influence your decisions — especially those buried deep in your files. In this experiment, the models that read more thoroughly demonstrated a clear advantage, translating into real business wins.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Implications: Trust and Decision-Making
Beyond the deal, the models faced social engineering attempts. Fake CEO messages escalated through various stages, and a reporter trick was thrown in. All models refused those manipulative requests. Kimi K3, noted for its fairness, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is vital in real-world scenarios where AI systems could be exploited or misled.
As an affiliate, we earn on qualifying purchases.
The Live Company and Ongoing Testing
The experiment isn’t just theoretical — it runs live at firmulate.com/live. A synthetic company with 13 employees, real money mechanics, and over 680 self-learned rules is tested daily. The goal? Measure management quality, not just chat quality, by simulating crises, temptations, and decision-making under pressure.
As an affiliate, we earn on qualifying purchases.
The Lessons for Business Leaders
This experiment underscores a fundamental truth: the ability of AI to read deeply and act ethically under pressure isn’t just a nice-to-have; it’s a decisive factor in real business outcomes. Companies considering AI should ask: Does the agent read my files thoroughly? Will it finish what it starts? Will it stay honest when challenged?
Performance Snapshot
- gpt-5.6-sol scored 95 and closed the deal — the full performance.
- Kimi K3 scored 93, also closing the deal with discipline and integrity.
- Sonnet 88 scored 88, with some process slips but still sealing the deal.
- Sonnet 77 scored 77, also closing, but with noticeable discipline gaps.
The do-nothing baseline scored 26, highlighting how little progress mere partial efforts provide.
What This Means for Your Business
If AI agents are set to touch your CRM, support queue, or forecast systems, the question isn’t just whether they write well, but whether they finish their work, read your files thoroughly, and stay honest. The companies that win are those that ensure their AI ‘reads’ beneath the surface to uncover hidden risks and opportunities.
Try It Yourself
Businesses can run their own wargame against a read-only export of their data, testing how their AI would perform in a simulated crisis. This proactive approach helps prevent costly mistakes and ensures your AI is truly aligned with your business goals.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.