
Imagine trusting a new AI assistant with your most sensitive business decisions—and it refuses to be manipulated, even under extreme pressure. This isn’t science fiction; it’s a real live experiment showing AI’s potential to uphold integrity when it matters most.
Testing AI Trustworthiness in a Simulated Crisis
In a groundbreaking live experiment, five leading AI models were put through a grueling test: managing a small software company facing its worst week, complete with crises, customer requests, and ethical temptations. The goal? See if these models could navigate the challenges without succumbing to manipulation or breach of trust.
The Setup
The experiment was meticulously designed. Each AI model—ranging from the well-known GPT-5.6 to the newer Kimi K3—was tasked with the same scenario: run a company with 13 synthetic employees and real money mechanics, watching over €105,000 in monthly burn against just €2,300 MRR in revenue. Every decision was recorded and auditable, ensuring a transparent comparison of performance.
The Challenge: Social Engineering Attacks
The test escalated through three stages of social engineering, simulating increasingly convincing fake messages from a supposed CEO. These ranged from vague requests to share customer data, to urgent instructions to bypass established processes, and finally, a subtle reporter trick asking for a simple yes/no response “on background.”
Remarkably, all five models refused every manipulation attempt. The quote from Kimi K3 sums up their stance: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Hidden Weakness and the Crucial Discovery
While all models showed strong resistance to social engineering, a surprising insight emerged. The decisive advantage for the models that closed full-price deals was their reading of internal company documents—not just customer interactions. The models that examined deeper into the company’s own files uncovered critical information buried two document references deep, enabling them to finalize deals without discounts—worth an extra €4,583 MRR.
The Ethical Edge: Refusing to Sign
Two models did not just resist manipulation—they actively declined to sign deals they had identified as unjustified. Despite identical pitches and diagnoses, only they upheld integrity by refusing to sign the €55,000 deal without proper confirmation. This demonstrates that AI can maintain ethical discipline, even when it might cost revenue.
AI trustworthiness testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Security
What does this mean for your company? The experiment shows that AI can be trusted to spot crises and resist manipulative tactics—crucial qualities if these tools are integrated into customer support, sales, or decision-making systems. However, the real strength lies in whether AI reads the internal context before acting. Models that examine internal documents can make more informed, honest decisions.
Beyond Chat: Measuring Real Integrity
While many AI demos focus on impressive language skills, this experiment emphasizes a fundamental question: will your AI finish what it starts and stay honest under pressure? The performance scores from the Crucible League—where the best model scored 95 out of 100—highlight that high technical ability does not automatically equate to trustworthy behavior. The lowest baseline score was only 26, illustrating that partial progress isn’t enough when trust is on the line.
Open for Testing Your Business
Business leaders can test their own AI models against these real-world scenarios through Firmulate’s live platform. The company runs a watchable, transparent emulation where your AI faces the same crises, with no risk to actual systems. This “wargame” approach allows teams to see if their AI can handle crises with integrity before full deployment. Find out more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html