firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI assistant that, despite doing almost nothing, still earns a baseline score of 26 out of 100. Sounds strange? In the world of AI benchmarking, this “do-nothing” score reveals surprising truths about trust, progress, and reliability — lessons that matter even for your home decor projects or gift shops.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Values Honesty and Progress

At a recent live experiment conducted by Firmulate, four AI models were tasked with running a simulated small software company through its worst week. The scenario involved real crises, customer demands, and ethical tests, all designed to see how AI agents would perform under pressure. The key takeaway? Every AI spotted every crisis and refused manipulative or dishonest requests. Yet, only two of them completed the job by signing a €55,000 deal.

What’s intriguing is that all four models scored at least 26 points—even the one that did nothing but stay honest. This baseline score, called the ‘do-nothing’ score, is not zero. Why? Because even in the absence of action, the AI’s ability to recognize problems and refuse bad behavior counts for something. Partial progress, such as catching a crucial document or resisting a scam, adds to the score. Conversely, a single breach of trust caps the total grade—no amount of good work can compensate for dishonesty.

Amazon

AI ethics and trust software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Deeply Matters

The experiment uncovered an important insight: the decisive advantage for some models came from reading a specific document deep inside the company’s files. The models that read and understood this buried information secured the deal at full price, adding €4,583 in monthly recurring revenue. This shows that in complex environments, even a small detail can make or break the outcome.

How Trust and Discipline Are Tested

Beyond handling crises, the models faced social engineering attempts, including staged fake CEO messages and reporter tricks. All refused these manipulative tactics, indicating a strong understanding of ethical boundaries. Kimi K3, one of the top performers, explained its refusal by saying, “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet, discipline slipped with Opus 4.8, the most thorough participant, which left a potential deal on the table by failing to escalate an issue properly. This highlights that even the most advanced AI can falter in process discipline, especially when overwhelmed or when rules are not enforced strictly.

The Reality of Running a Live Business with AI

The experiment isn’t just theoretical. It’s run on a real, live simulated company with 13 synthetic employees managing actual money — burning €105,000 each month against a revenue of only €2,300. Every decision is versioned daily, and over 680 self-learned rules guide its operations. The entire setup is transparent and accessible at firmulate.com/live.

This real-world context underscores that AI’s true value isn’t measured by how well it chats or answers questions but by whether it can finish tasks reliably, read crucial information, and stay honest under duress. The benchmark’s live, auditable environment mimics the pressures your business faces—whether in home decor sales, gift packaging, or customer service.

What This Means for Your Business

So, why should you care about a benchmark that includes a ‘do-nothing’ score or emphasizes trust? Because the core question businesses should ask their AI providers is: Will your AI finish what it starts, read the critical details, and stay honest when the pressure is on? An AI that scores low on these aspects, regardless of how well it mimics human conversation, isn’t reliable for managing real-world operations.

Remember, the current league table shows that GPT-5.6 scores 95, with the ability to find buried facts and close deals. Kimi K3 scores 93, demonstrating excellent discipline and honesty. These high performers can be trusted to handle your CRM, support queues, or even forecast demand without slipping. Meanwhile, the lower scores remind us that even the best AI can fail if it’s not carefully trained and tested.

How You Can Prepare — and Wargame Your AI Workforce

Firmulate offers enterprises a way to test their AI models against similar scenarios through a ‘wargame’—a sandbox environment where your AI can face real crises without any risk to your actual business. This approach helps you identify weaknesses before deployment, ensuring your AI behaves ethically and effectively in high-pressure situations. Details are available at firmulate.com/pilot.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Robot Vacuum Reality Check: What Floor Types They Handle Best

Perfectly suited for hard floors and low-pile carpets, robot vacuums face challenges with thicker surfaces—discover how to maximize their cleaning potential!

Smart Thermostats: The Comfort Settings Older Adults Tend to Prefer

The comfort settings older adults prefer in smart thermostats automatically learn routines and offer convenient control, but there’s more to discover about enhancing your home.

The Smart Device Compatibility Check That Saves Major Frustration

Finding the right smart device compatibility check can prevent major frustration—continue reading to discover how to connect everything seamlessly.

Charging Stations at the Bedside: The Cleaner, Safer Setup

Here’s a safe and organized way to set up your bedside charging station that might just change your nightly routine—find out more.