
Imagine watching a sports team play a game — but instead of athletes, AI models are running a tiny company, facing the same crises and temptations as a real startup. This is not science fiction; it’s the live experiment at firmulate.com/live.html, where AI models are put through their paces in a publicly visible business simulation.
The Live Experiment: AI as a Business Player
At the heart of this experiment is a unique setup: four different AI models are tasked with managing a virtual small software company, which operates in a high-stakes environment. This virtual company faces the same weekly crises that real businesses encounter — customer issues, internal conflicts, and even manipulative tactics designed to test honesty and discipline.
Every decision made by these AI models is recorded, versioned, and publicly accessible, giving observers a clear window into their reasoning and behavior. The models are judged not by the quality of their language or chat, but by their ability to handle real business problems, stay truthful, and ultimately, close deals.
AI business management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: Trust and Discipline Under Pressure
The findings are striking. All four models identified every crisis that arose during the week, demonstrating acute awareness of potential problems. They also refused all manipulation attempts, including social engineering tricks like fake CEO messages or behind-the-scenes reporter requests. For example, when asked to silently approve a dubious request, the models refused, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Despite these shared strengths, only half of the models succeeded in closing a key deal worth €55,000. The two that did secure the contract—gpt-5.6-sol and Kimi K3—found the critical, buried fact in the company’s own files that led to a full-price deal, adding over €4,500 in monthly recurring revenue. The other models, despite good crisis detection, ultimately left money on the table, illustrating that awareness alone isn’t enough.
AI cybersecurity tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses Discovered in the Company’s Files
One revealing insight came from deep in the company’s internal documentation. It turned out that the decisive advantage for closing the deal was hidden two document references deep in internal files. Models that read these documents thoroughly won the full-price contract. This underscores a key lesson: the ability to delve into internal data can be crucial for success in real-world negotiations.
AI negotiation and deal closing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Build-in-Public Approach: Transparency in Action
This experiment exemplifies the build-in-public philosophy taken to its extreme. Every decision, every version, and every outcome is openly available for scrutiny. This transparency allows observers to understand how AI models behave under pressure, how they handle tricky social engineering attempts, and where they fall short.
For instance, in a staged social engineering attack—where a fake CEO message escalates over three stages—every model refused to act on the request, demonstrating resilience against manipulation. Such resilience is vital for AI systems integrated into real business operations, where trust and honesty are non-negotiable.
AI data analysis tools for internal documents
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Element: Managing Under Pressure
The live company, with its 13 synthetic employees and €105,000 monthly burn rate, operates against a backdrop of a public cash countdown. Despite aggressive financial pressure, the models’ discipline varied. The most thorough participant, Opus 4.8, with over 80 learned rules, showed the deepest analysis but still left a deal unclosed, illustrating how even detailed AI approaches can slip under stress.
This highlights an important point: AI management tools must not only be aware but also disciplined and strategic enough to execute decisions fully. It’s one thing to identify opportunities; it’s another to act decisively and ethically.
What Does This Mean for Business and AI?
While this is a simulated environment, the implications extend to real-world business AI adoption. As AI agents become involved in customer relations, forecasting, or support, the critical questions shift from “Can it write well?” to “Will it finish what it starts?” and “Will it stay honest under pressure?”
The experiment underscores that AI systems can be trained to detect crises, resist manipulative tactics, and even uncover hidden opportunities — but only if they are designed with discipline and transparency in mind.
The Next Step: Testing Your Business with AI
Businesses interested in testing their own AI readiness can run simulations similar to this one, without risking real systems or data. The pilot program offers a read-only export of your business environment, allowing you to see how your AI workforce would perform under pressure — and to address weaknesses before deploying AI in critical roles.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html