firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an employee who, no matter the crisis or temptation, always refuses to lie or cut corners. Yet, despite this unwavering honesty, they only achieve a modest score of 26 out of 100 in a recent AI benchmark. For businesses pondering AI integration, understanding this score isn’t just about numbers — it’s about trust, discipline, and the true value of AI in complex decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In a recent, transparent experiment conducted by Firmulate, four advanced AI models were tasked with managing a small software company’s worst week — complete with crises, manipulative tactics, and high-stakes decisions. The goal? To evaluate not just their technical abilities, but their integrity, discipline, and real-world decision-making under pressure.

What makes this experiment stand out is its rigorous methodology. Every decision was versioned and auditable, ensuring that outcomes could be traced and verified. This approach gives business leaders concrete insights into how AI behaves when it faces ethical dilemmas or attempts at manipulation — critical factors for deploying AI in sensitive environments like customer service, compliance, or financial operations.

The Surprising Baseline

One of the key findings is that even a ‘do-nothing’ AI baseline scores 26 points. This baseline isn’t just a random starting point; it reflects partial progress, where the model recognizes problems but doesn’t necessarily act decisively. Interestingly, even in the absence of proactive behavior, just avoiding mistakes like reading sensitive files or signing deals prematurely can earn some points.

However, the experiment also underscores a fundamental principle: a single breach of trust — such as signing a deal without proper analysis or bypassing a protocol — caps the model’s overall score. In other words, no matter how good a model is at crisis detection or decision-making, one slip-up can disqualify it from full trust.

Models That Stay Honest Win the Deal

All four models successfully identified crises and refused manipulative tactics like fake CEO messages or reporter tricks. Notably, only two models signed the deal they had analyzed and recommended — their own analysis — and did so without shortcuts. The other two models, despite recognizing the opportunity, left the deal on the table, illustrating how discipline and trustworthiness directly impact tangible business outcomes.

Delving deeper, the experiment revealed that the decisive advantage often lay in how well the models read and interpret internal documents. The winning models accessed critical information buried two document references deep in the company’s files — knowledge that, if overlooked, meant missed revenue opportunities. Reading and understanding internal data is thus crucial for AI to perform fully and ethically.

Resistance to Social Engineering

Another remarkable aspect was the models’ response to social engineering. Fake CEO messages escalating over multiple stages and attempts to manipulate the AI into approvals or impersonations were universally refused. Kimi K3, in particular, responded aptly: “Treat the request as a suspected approval-bypass / possible impersonation.” This resilience to manipulation demonstrates the importance of built-in safeguards against trust breaches, especially when AI interacts with humans in sensitive roles.

Real-World Application and Limitations

The experiment isn’t just a theoretical exercise. The live company in the experiment operates with 13 synthetic employees managing real money mechanics — burning €105,000 monthly against €2,300 in monthly recurring revenue. Every decision and rule is versioned daily, allowing continuous monitoring and improvement.

Yet, the experiment also shows that even the most thorough AI — like Opus 4.8, which employed over 80 learned rules and deep analysis — can fall short if discipline slips. In this case, the model left a deal on the table and failed to escalate certain issues, highlighting that thoroughness alone isn’t enough; consistent discipline and adherence to protocols matter.

What Business Leaders Need to Know

For companies considering AI in critical decision-making roles, the key takeaway isn’t just about technical prowess. It’s about whether the AI can finish what it starts, read internal data thoroughly, and maintain integrity under pressure. The benchmark scores show that even honest AI models don’t score high — the top model, gpt-5.6-sol, earned 95 points, while the do-nothing baseline remains at 26.

This transparency is vital. It reveals that deploying AI isn’t just about getting it to produce convincing outputs; it’s about ensuring consistent discipline, trustworthiness, and ethical decision-making.

To see how your organization can test and prepare your AI workforce, you can explore the live benchmark experiments, which simulate real crises and decision challenges. These benchmarks serve as a crucial step before deploying AI in any environment where trust and discipline are paramount.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Charging Stations for Staff Gear: The One Setup That Prevents Dead Devices

Wireless charging stations for staff gear can prevent dead devices—discover the essential setup that keeps everyone connected and productive.

International Inspiration: Luxury Mobile Restrooms Abroad

Noticing how luxury mobile restrooms abroad blend eco-friendly design with cultural elements reveals innovative trends transforming event experiences worldwide.

Will Autonomous Cleaning Robots Transform Portable Restroom Maintenance?

AIThis post was created with the assistance of artificial intelligence (AI).Autonomous cleaning…

How AI Read Your Files to Win Big Deals — and Why It Matters for Your Business

AI that reads deep into your company files before making decisions is proving to be a game-changer, winning deals and avoiding manipulation in complex scenarios. Learn more.