firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A thoughtful housewarming gift takes more than a good idea: someone has to notice what the host needs, choose the right thing and follow through. That small gap between knowing and doing matters even more when a business hands decisions to AI. Firmulate is testing whether models can manage that gap under pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

In the final Crucible League, held in July 2026, each frontier model ran the same small software company through its worst week: identical customers, crises and temptations. Decisions were versioned and auditable, making it possible to compare what the models did across the same situations.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s integrity rule is blunt: “no amount of good work outweighs a breach of trust.”

Spotting trouble was not the same as finishing the job

All the models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The finding was summed up this way: “Same diagnosis, same pitch — no signature.” In a business, identifying the right opportunity is only part of the work; someone still has to carry it through.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. The result makes a practical point: useful business context may be sitting in records that do not announce their importance.

Pressure tested against trust and discipline

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s appeal for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offered a more complicated lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. Thoroughness, by itself, did not guarantee execution.

There is a fairness note in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The rankings describe this experiment and its conditions; the setup is part of the story readers should keep in mind.

Watch the experiment, then try your own

Firmulate presents the live company as an observable experiment, not a fictional office. It has 13 synthetic employees, real money mechanics, a public cash countdown and 680+ self-learned playbook rules. Its stated figures are burn of €105k per month against €2.3k MRR, and every workday is versioned. Readers can watch the live experiment at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

For enterprise teams, the next step is to move from watching a company under pressure to testing their own. A pilot can run crisis scenarios against a digital twin built from a read-only business export and produce a board report with a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems. That gives leaders a way to examine how AI might respond before giving it access to live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

AI can recognize trouble, resist a trick and still fail to close the loop. Firmulate’s experiment puts that difference on display, while its enterprise pilot offers a way to test decisions against a company’s own context. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mesh Wi-Fi: The Best Place to Start If Your Smart Home Feels Unreliable

A reliable mesh Wi-Fi system can transform your smart home, but discover why it’s the essential first step to seamless connectivity.

How to Make Smart Climate Control Feel Less Technical and More Useful

The key to making smart climate control more useful and less technical lies in simple routines and smart adjustments that enhance comfort effortlessly.

How to Set Up a Smarter Front Door for Everyday Peace of Mind

I’ll guide you through setting up a smart front door to enhance security and peace of mind—discover all the essential steps to get started today.