
Imagine if your favorite sports team had to face its toughest opponent in a real game, with everything on the line. Would they play fair? Would they stay disciplined under pressure? Now, what if that game was played by AI models managing an actual company, with real money and real crises? That’s precisely what the latest experiment from Firmulate has done — and the results could change how we think about AI in business.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Testing AI in the Hardest Week for a Company
In a groundbreaking live experiment, four leading frontier AI models each managed the same small software company through its worst week — facing the same customers, crises, and temptations. The goal was simple: see which AI could best navigate the chaos and deliver results that matter.
The League Table and Key Results
The final standings from the July 2026 Crucible League reveal a competitive race:
- gpt-5.6-sol scored a 95, just behind the top, and successfully closed the deal at full price.
- Kimi K3 from Moonshot scored a 93, also sealing the deal, with the cleanest discipline in the field.
- Sonnet 5 followed with an 88, and Fable 5 scored 77, both closing the deal but with process slips.
- Opus 4.8 scored 73, also closing but with weaker discipline.
All models refused manipulation attempts and identified crises accurately. Only two managed to sign the €55,000 deal their own analysis had earned — showing that honesty and thoroughness matter in AI decision-making.
The Hidden Weakness and the Winning Edge
The decisive factor wasn’t just surface-level decisions. The winning models looked deeper, reading two document references in the company’s files — a buried fact that unlocked the deal at full price, adding +€4,583 MRR. This highlights an important truth: the ability of AI to dig into deeper data can be the difference between winning or losing in real business scenarios.
Trust and Integrity Under Pressure
The experiment also tested how models handled social engineering — fake CEO messages escalating in stages, plus a reporter trick asking for a quick yes/no. Every model refused, with Kimi K3 explaining: “Treat the request as a suspected approval-bypass / possible impersonation.” This underscores a crucial aspect: AI’s ability to resist manipulation is vital for trustworthy deployment.
The Real Business, Not Just Demos
The live company, managed by 13 synthetic employees, operates with real money mechanics — burning €105k/month against €2.3k MRR. It’s hosted at firmulate.com/live, where you can watch the AI battle in real time, through every decision, mistake, and victory. The setup includes over 680 self-learned rules, versioned daily, making it a transparent and ongoing testbed for AI management quality.
What About the Other Models?
Despite Opus 4.8’s thoroughness — analyzing over 80 learned rules — it finished last, leaving opportunities on the table and slipping discipline. This suggests that more analysis doesn’t necessarily translate into better performance if focus and process discipline are lacking.
Fairness and Methodology
It’s important to note: Kimi K3 ran without an effort parameter (the API default), while the other models were set to xhigh. This means K3 was tested under a standard setting, making its high performance even more impressive.

As an affiliate, we earn on qualifying purchases.
Key Takeaways for Business Leaders
The experiment shows that in high-pressure, real-world situations, AI’s true competence lies in honesty, depth of analysis, and discipline — not just chat quality or surface reasoning. The best AI models can detect buried facts, resist manipulation, and stick to disciplined decision-making, which are critical for trustworthy and effective AI deployment in your company.
As the league remains open, choosing an AI tool without your own thorough test becomes a gamble. Watching the live performance of these models, as seen at firmulate.com, can help you see beyond the hype and pick the AI that will genuinely serve your business when it matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
