
Imagine a company that operates entirely without human employees, yet still faces the harsh realities of cash flow, customer crises, and tough decisions—every single day. Sound impossible? Welcome to the frontier of artificial intelligence-driven management, where a real, live company is testing just how far AI can go in running a business.
The Live Company: A Living Experiment in AI Management
At first glance, it looks like any small software firm: a handful of staff, a steady revenue stream of €2,300 per month, and a looming cash deficit of €105,000 each month. But this company is unlike any other—it’s entirely virtual, composed of 13 synthetic employees powered by cutting-edge AI models. Every workday, its decisions, crises, and strategies are recorded, versioned, and made public at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
Testing the Limits of AI Decision-Making
The experiment pits four leading AI models against each other by running the same challenging week in the company’s life—one filled with customer complaints, urgent crises, manipulative tactics, and tricky negotiations. These models include GPT-5.6-SOL, Kimi K3, Sonnet 5, and Opus 4.8, each with unique strengths and flaws. Their task: navigate the worst week possible, keep the company afloat, and close deals where possible.
Key Findings from the Experiment
- All four models identified every crisis the company faced, from technical failures to customer blowups.
- Every model refused to fall for manipulative tricks, such as fake CEO messages or reporter ‘background’ requests—maintaining integrity under pressure.
- Only two models managed to close the €55,000 deal their own analysis had earned. The same diagnosis and pitch, but only these two signed the contract.
- Deep in the company’s own files—two documents down—they uncovered a critical piece of information needed to win the deal, which others missed. The model that read the file won the deal at a full €4,583 in monthly recurring revenue.
The Surprising Weaknesses and Lessons
The most thorough participant, Opus 4.8, with over 80 learned rules and the deepest analysis, ultimately finished last in the deal. Its discipline waned, and it failed to escalate issues appropriately, leaving potential revenue on the table. All models, despite their different approaches, showed vulnerabilities rooted in process discipline when under duress.
The Broader Implications for Business and AI
This experiment isn’t just about running a virtual company; it highlights critical questions for any business considering AI automation:
- Will AI decision-makers read and interpret all necessary information—especially hidden, buried facts—before acting?
- Can AI maintain honesty and integrity when faced with manipulation or social engineering tactics?
- Are AI systems capable of closing deals and completing useful work, not just generating convincing dialogue?
While this company operates in the realm of software, the implications extend to sectors like customer support, sales, and operations—areas where AI’s ability to finish what it starts, stay honest under pressure, and read deeply into data is crucial.
Watching the Future Unfold
Watch this ongoing experiment at firmulate.com/live. Every weekday, the company’s decisions are recorded, versioned, and displayed for all to see—an unprecedented build-in-public effort to understand how AI can truly manage a business. The results so far are clear: AI can handle crises and resist manipulation, but execution discipline and strategic depth are still evolving.

This live experiment offers a rare glimpse into AI’s potential—and its limitations—in real-world business management. As AI models compete and improve, their ability to finish tasks, read critical data, and act honestly will determine whether AI becomes a trustworthy partner in running our companies or remains a powerful but imperfect tool.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html