
Imagine a sports team facing its toughest week yet—crises, tough decisions, and the pressure to perform. Now, substitute the team with AI models managing a small software company under real-world stress. Which AI would come out on top? The answer isn’t just about speed or clever talk; it’s about integrity, discipline, and results. That’s the core of a groundbreaking live experiment by Firmulate, where AI models are put through a management gauntlet that mimics a real company’s worst week.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
How Firms Are Testing AI in the Real World
In an unprecedented experiment, four frontier AI models are running a real small software company during its most challenging week. This isn’t staged or simulated; every crisis, customer interaction, and temptation to cheat is identical across all models. The goal? To see which AI can handle management tasks with integrity, discipline, and effectiveness.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Crisis, Different Minds
Each model was tasked with making decisions, from addressing customer complaints to negotiating deals. The company’s situation was dire, losing money every day with a cash runway that was quickly shrinking. The models had access to the same data, the same company history, and faced the same external pressures. Every decision was recorded and auditable, providing a transparent window into how each AI behaved under stress.
business AI chatbot for customer service
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Key Findings: Integrity and Results Matter
All four models successfully identified crises and refused manipulative attempts such as social engineering scams. They faced escalating fake CEO messages and a reporter trick that asked for a simple yes/no answer on background—each refused to be duped. But the true test was whether they could close a critical deal worth €55,000. Only two models managed to sign the deal, and intriguingly, both had discovered a crucial internal document reference that their competitors missed.
The Hidden Weakness
The decisive advantage went to the models that read deeper into the company’s files. They identified and leveraged a buried fact that gave them the full deal at a higher monthly recurring revenue (+€4,583 MRR). Conversely, models that failed to delve into the document references left the deal on the table, missing out on significant gain.
As an affiliate, we earn on qualifying purchases.
The Personalities of AI Managers
One model, Opus 4.8, was the most thorough, analyzing over 80 learned rules and providing detailed, cautious reasoning. Yet, it was the last to close, leaving the deal unsealed and slipping discipline. Kimi K3 took a different approach—more direct and succinct, and notably, it ran without an effort parameter, meaning it was operating at a default, high-effort level. The other models, which ran at an even higher effort setting, showed slightly different behaviors but ultimately refused to cheat or manipulate.
As an affiliate, we earn on qualifying purchases.
Beyond the Demos: Real Company, Real Money
The experiment is conducted within a live, functioning company environment that handles real money mechanics—burning €105,000 monthly against a revenue of just €2,300. The company’s operations are continuously versioned and observable at firmulate.com/live. This ongoing experiment illustrates how these models perform not just in tests but in a real business context with actual financial stakes.
Why This Matters for Business Leaders
The core lesson from Firmulate’s experiment isn’t about how well AI can chat or mimic human conversation. It’s about whether AI can finish what it starts, read critical information before acting, and stay honest under pressure. For businesses integrating AI, this experiment highlights that the real measure of a model isn’t just in its cleverness—it’s in its discipline and integrity.
The Top Performers and the League Table
- gpt-5.6-sol: Scored the highest with 95 points, found the buried fact, and closed the deal successfully.
- Kimi K3: Achieved a close score of 93, operating without an effort parameter, yet still closing the deal with the cleanest discipline.
- Sonnet 5: Scored 88, closed the deal but with a few process slips.
- Fable 5: Scored 77, also closed the deal but showed more slips in process discipline.
See It Live and Decide for Yourself
This isn’t just a demo. The company’s entire operation runs every business day, watched by real people and real money. You can see the decisions unfold, read the actual conversations, and even run the same management wargame yourself at firmulate.com/quiz.html.

AI models can be tested in real-world management scenarios to see if they stay disciplined and honest under pressure. Results show that integrity and thoroughness can make the difference in closing crucial deals—vital insights for businesses considering AI-driven decision-making.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
