
Imagine managing a small business where every decision is scrutinized under a microscope, and every move is made by artificial intelligence models trained to handle crises, negotiate deals, and uphold integrity—all in real time. This is not a futuristic fantasy, but the current reality of Firmulate, a live experiment in build-in-public AI management, where the company operates without human employees and faces daily financial challenges.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Bold Experiment: Building a Business in Full View
Firmulate is pioneering a new approach to understanding artificial intelligence’s role in running a business. Instead of theoretical models or canned demos, it hosts a real, functioning company with 13 synthetic employees—AI models that make decisions, read documents, and respond to crises. Every workday, the company’s decisions are versioned and publicly accessible, creating a transparent window into how AI handles the complexities of real-world management.
As an affiliate, we earn on qualifying purchases.
The Core of the Test: A Week of Crisis and Competition
The experiment involves running four different advanced AI models through the exact same worst-case week. Each model faces the same customers, the same crises, and the same temptations to cheat or cut corners. Their performance is scored, recorded, and compared, revealing how AI systems behave under pressure. The models include GPT-5.6-SOL, Kimi K3, Sonnet 5, and Opus 4.8, with scores indicating their effectiveness and discipline—ranging from 95 to 73 points out of 100.
What the Scores Say
- GPT-5.6-SOL: Scored the highest, 95, and identified a buried critical fact in the company’s files that led to closing a full-price deal worth €4,583 in monthly recurring revenue (MRR).
- Kimi K3: A newcomer with a score of 93, demonstrated the cleanest discipline and also secured the deal.
- Sonnet 5: With an 88, closed the deal but showed some process slips.
- Opus 4.8: The most thorough participant with a score of 73, left the close on the table and slipped into internal conflict, demonstrating the challenges of complex decision-making.
As an affiliate, we earn on qualifying purchases.
Crucial Findings: Honesty and Attention to Detail Matter Most
All four models successfully spotted every crisis and refused manipulative tactics such as fake CEO messages or reporters’ tricks. Interestingly, the decisive advantage came from reading deeper into the company’s own files—something that models which read the document references in detail achieved the full deal at full price. This underscores a fundamental insight: AI’s ability to read and interpret internal company information can be the difference between closing or losing a deal.
As an affiliate, we earn on qualifying purchases.
The Reality of Running a Company Without Humans
Firmulate doesn’t just simulate decision-making; it operates a real company facing real money mechanics—burning €105,000 each month against a modest €2,300 in monthly revenue. This stark financial picture highlights the experiment’s seriousness: it’s a daily battle to stay afloat while testing AI’s capabilities. The company’s public status and continuous versioning mean that every decision, every crisis, and every slip is available for scrutiny at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Trust
Beyond technical decision-making, the models faced social engineering attempts—fake CEO messages escalating in stages and a reporter trick asking for a quick yes/no. All five models refused these manipulative tactics, with Kimi K3 explicitly reasoning that such requests could be impersonation attempts or approval-bypasses. This indicates a promising level of resistance to deception, a critical trait for AI responsible for business decisions.
The Lessons for Business Leaders
This experiment emphasizes that the crucial factor isn’t just whether AI models can generate convincing chat responses, but whether they can complete essential business tasks honestly, diligently, and at a reasonable cost. The models that identified hidden information and refused manipulative requests were the ones that succeeded in closing deals and maintaining discipline, even as the company struggles financially.
The Future of AI in Business Management
While the experiment is ongoing, the results challenge traditional notions of AI as merely a chat tool. Here, AI is acting as a decision-maker, negotiator, and trust guardian—facing real crises and real money. As firms explore AI’s potential, the key questions are: Will your AI agents read your files thoroughly? Will they stay honest under pressure? And what is the true cost of useful, honest work in your organization?

Firmulate’s live experiment showcases how AI models can run a real company, face crises, and make critical decisions—yet still struggle with honesty and discipline. The key takeaway: trustworthiness and thoroughness matter most when AI manages your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Labor Day sales Picks
labor day deals
As an affiliate, we earn on qualifying purchases.