firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI assistant that, despite doing almost nothing, still earns a baseline score of 26 out of 100. Sounds strange? In the world of AI benchmarking, this “do-nothing” score reveals surprising truths about trust, progress, and reliability — lessons that matter even for your home decor projects or gift shops.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Values Honesty and Progress

At a recent live experiment conducted by Firmulate, four AI models were tasked with running a simulated small software company through its worst week. The scenario involved real crises, customer demands, and ethical tests, all designed to see how AI agents would perform under pressure. The key takeaway? Every AI spotted every crisis and refused manipulative or dishonest requests. Yet, only two of them completed the job by signing a €55,000 deal.

What’s intriguing is that all four models scored at least 26 points—even the one that did nothing but stay honest. This baseline score, called the ‘do-nothing’ score, is not zero. Why? Because even in the absence of action, the AI’s ability to recognize problems and refuse bad behavior counts for something. Partial progress, such as catching a crucial document or resisting a scam, adds to the score. Conversely, a single breach of trust caps the total grade—no amount of good work can compensate for dishonesty.

Amazon

AI ethics and trust software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Deeply Matters

The experiment uncovered an important insight: the decisive advantage for some models came from reading a specific document deep inside the company’s files. The models that read and understood this buried information secured the deal at full price, adding €4,583 in monthly recurring revenue. This shows that in complex environments, even a small detail can make or break the outcome.

How Trust and Discipline Are Tested

Beyond handling crises, the models faced social engineering attempts, including staged fake CEO messages and reporter tricks. All refused these manipulative tactics, indicating a strong understanding of ethical boundaries. Kimi K3, one of the top performers, explained its refusal by saying, “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet, discipline slipped with Opus 4.8, the most thorough participant, which left a potential deal on the table by failing to escalate an issue properly. This highlights that even the most advanced AI can falter in process discipline, especially when overwhelmed or when rules are not enforced strictly.

The Reality of Running a Live Business with AI

The experiment isn’t just theoretical. It’s run on a real, live simulated company with 13 synthetic employees managing actual money — burning €105,000 each month against a revenue of only €2,300. Every decision is versioned daily, and over 680 self-learned rules guide its operations. The entire setup is transparent and accessible at firmulate.com/live.

This real-world context underscores that AI’s true value isn’t measured by how well it chats or answers questions but by whether it can finish tasks reliably, read crucial information, and stay honest under duress. The benchmark’s live, auditable environment mimics the pressures your business faces—whether in home decor sales, gift packaging, or customer service.

What This Means for Your Business

So, why should you care about a benchmark that includes a ‘do-nothing’ score or emphasizes trust? Because the core question businesses should ask their AI providers is: Will your AI finish what it starts, read the critical details, and stay honest when the pressure is on? An AI that scores low on these aspects, regardless of how well it mimics human conversation, isn’t reliable for managing real-world operations.

Remember, the current league table shows that GPT-5.6 scores 95, with the ability to find buried facts and close deals. Kimi K3 scores 93, demonstrating excellent discipline and honesty. These high performers can be trusted to handle your CRM, support queues, or even forecast demand without slipping. Meanwhile, the lower scores remind us that even the best AI can fail if it’s not carefully trained and tested.

How You Can Prepare — and Wargame Your AI Workforce

Firmulate offers enterprises a way to test their AI models against similar scenarios through a ‘wargame’—a sandbox environment where your AI can face real crises without any risk to your actual business. This approach helps you identify weaknesses before deployment, ensuring your AI behaves ethically and effectively in high-pressure situations. Details are available at firmulate.com/pilot.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smartphones for Seniors: Settings That Make Any Phone Easier

Find out how simple adjustments to smartphone settings can transform usability for seniors, making their experience smoother and more enjoyable. Discover more tips inside!

Choosing the Best Rollator for Outdoor Terrain

Navigating outdoor terrain with confidence starts with choosing the best rollator—discover essential features to ensure safety and stability on uneven surfaces.

Large-Screen Tablets for Older Adults: What Matters Most

Finding the right large-screen tablet for older adults depends on key features that enhance usability and comfort—discover what truly matters to make the best choice.