
Imagine a company that operates entirely without human employees, making tough decisions in real time—yet it struggles to stay afloat financially. Now, picture watching this experiment unfold live, with every move scrutinized and every crisis tackled by artificial intelligence. Welcome to the world of Firmulate, where AI is not just assisting but running a company, and you can see it all happen in real time.
The Radical Concept of Firmulate
Firmulate is a one-of-a-kind experiment that pushes the boundaries of AI’s capabilities in management and decision-making. It features a virtual company with 13 synthetic employees, each guided by complex, self-learned rules—over 680 of them—that govern daily operations. The company’s financials are transparent and public: it burns €105,000 every month against a modest €2,300 in monthly recurring revenue (MRR). This is a real-time, open window into the struggles of a business fighting to survive.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI in Crisis Mode
The experiment involves running four advanced AI models through the same challenging scenario: a week filled with crises, customer pressures, and ethical temptations. Each model is given identical information, including critical details buried in internal files—not just surface-level customer data. This setup tests their ability to uncover hidden insights, make trustworthy decisions, and resist manipulation.
business ethics AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Outstanding Performance and Stark Limitations
All four models demonstrated exceptional crisis awareness, identifying every problem that arose. They also refused every attempt at manipulation—such as social engineering tricks and fake CEO messages—showing a strong adherence to ethical standards. However, only two managed to conclude the week by closing a significant deal valued at €55,000. Interestingly, the models that succeeded in sealing the deal were also the ones that read and understood the company’s internal documentation best, revealing a critical weakness in those that missed this step.
AI crisis management platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness – Reading the File
Despite their abilities, the models’ biggest vulnerability was a piece of buried internal information—hidden two document references deep—that, if discovered, could have unlocked the full potential of the business. The models that read the company’s internal files successfully identified this crucial fact and secured the deal at full price, adding over €4,500 in monthly recurring revenue.
AI enterprise decision support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Ethical Decision-Making
Beyond crisis management, the experiment tested social engineering resistance by attempting to influence the AI with staged CEO messages and a reporter trick. All five models refused these manipulative tactics, citing concerns over impersonation and bypassing approval protocols. Kimi K3, one of the models, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights the AI’s capacity to prioritize ethical safeguards under pressure.
The Real Company’s Financial Struggle
The live demonstration runs a virtual company on a tight budget—burning €105,000 monthly against just €2,300 in income. Every decision made by these AI models is versioned and auditable, with their decision-making rules constantly evolving. This transparency allows observers to understand how each choice impacts the company’s fragile survival, which is publicly documented at firmulate.com/live.html.
What This Means for the Future of Business Management
This experiment isn’t just a technical showcase. It raises vital questions about the role of AI in real-world management: Can AI agents reliably finish what they start? Will they stay honest under pressure? And how do we quantify the value of useful work when AI is the decision maker? As the leaderboard shows, while models like gpt-5.6-sol and Kimi K3 are close, there are still gaps—particularly in disciplined execution—even in AI systems designed for business.
Why Open Practice Matters
Firmulate’s approach of publicly running and versioning every decision offers a new standard in transparency. It invites business leaders and technologists to observe, learn, and prepare for AI-driven management. This build-in-public philosophy exposes both potential and pitfalls, emphasizing the importance of rigorous testing before trusting AI with critical operations.
Experience the Live Wargame
You can watch this ongoing battle for business survival at firmulate.com/live.html. See how AI models navigate crises, resist manipulation, and attempt to close deals—live and unfiltered. For those interested in testing their own systems or understanding AI’s decision-making under stress, Firmulate offers options to run similar simulations against your own business data—without risking your actual operations, at firmulate.com/pilot.html.

Watching AI run a virtual company in real time reveals its remarkable strengths—spotting crises, resisting manipulation, and closing deals—and its current limitations in disciplined execution. This build-in-public experiment offers a glimpse into how AI might soon shape trustworthy management, provided we understand and address its vulnerabilities.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html