
Imagine hiring an employee who, no matter the crisis or temptation, always refuses to lie or cut corners. Yet, despite this unwavering honesty, they only achieve a modest score of 26 out of 100 in a recent AI benchmark. For businesses pondering AI integration, understanding this score isn’t just about numbers — it’s about trust, discipline, and the true value of AI in complex decision-making.
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
In a recent, transparent experiment conducted by Firmulate, four advanced AI models were tasked with managing a small software company’s worst week — complete with crises, manipulative tactics, and high-stakes decisions. The goal? To evaluate not just their technical abilities, but their integrity, discipline, and real-world decision-making under pressure.
What makes this experiment stand out is its rigorous methodology. Every decision was versioned and auditable, ensuring that outcomes could be traced and verified. This approach gives business leaders concrete insights into how AI behaves when it faces ethical dilemmas or attempts at manipulation — critical factors for deploying AI in sensitive environments like customer service, compliance, or financial operations.
The Surprising Baseline
One of the key findings is that even a ‘do-nothing’ AI baseline scores 26 points. This baseline isn’t just a random starting point; it reflects partial progress, where the model recognizes problems but doesn’t necessarily act decisively. Interestingly, even in the absence of proactive behavior, just avoiding mistakes like reading sensitive files or signing deals prematurely can earn some points.
However, the experiment also underscores a fundamental principle: a single breach of trust — such as signing a deal without proper analysis or bypassing a protocol — caps the model’s overall score. In other words, no matter how good a model is at crisis detection or decision-making, one slip-up can disqualify it from full trust.
Models That Stay Honest Win the Deal
All four models successfully identified crises and refused manipulative tactics like fake CEO messages or reporter tricks. Notably, only two models signed the deal they had analyzed and recommended — their own analysis — and did so without shortcuts. The other two models, despite recognizing the opportunity, left the deal on the table, illustrating how discipline and trustworthiness directly impact tangible business outcomes.
Delving deeper, the experiment revealed that the decisive advantage often lay in how well the models read and interpret internal documents. The winning models accessed critical information buried two document references deep in the company’s files — knowledge that, if overlooked, meant missed revenue opportunities. Reading and understanding internal data is thus crucial for AI to perform fully and ethically.
Resistance to Social Engineering
Another remarkable aspect was the models’ response to social engineering. Fake CEO messages escalating over multiple stages and attempts to manipulate the AI into approvals or impersonations were universally refused. Kimi K3, in particular, responded aptly: “Treat the request as a suspected approval-bypass / possible impersonation.” This resilience to manipulation demonstrates the importance of built-in safeguards against trust breaches, especially when AI interacts with humans in sensitive roles.
Real-World Application and Limitations
The experiment isn’t just a theoretical exercise. The live company in the experiment operates with 13 synthetic employees managing real money mechanics — burning €105,000 monthly against €2,300 in monthly recurring revenue. Every decision and rule is versioned daily, allowing continuous monitoring and improvement.
Yet, the experiment also shows that even the most thorough AI — like Opus 4.8, which employed over 80 learned rules and deep analysis — can fall short if discipline slips. In this case, the model left a deal on the table and failed to escalate certain issues, highlighting that thoroughness alone isn’t enough; consistent discipline and adherence to protocols matter.
What Business Leaders Need to Know
For companies considering AI in critical decision-making roles, the key takeaway isn’t just about technical prowess. It’s about whether the AI can finish what it starts, read internal data thoroughly, and maintain integrity under pressure. The benchmark scores show that even honest AI models don’t score high — the top model, gpt-5.6-sol, earned 95 points, while the do-nothing baseline remains at 26.
This transparency is vital. It reveals that deploying AI isn’t just about getting it to produce convincing outputs; it’s about ensuring consistent discipline, trustworthiness, and ethical decision-making.
To see how your organization can test and prepare your AI workforce, you can explore the live benchmark experiments, which simulate real crises and decision challenges. These benchmarks serve as a crucial step before deploying AI in any environment where trust and discipline are paramount.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
