
A pool company’s worst week can start with a heat wave, a supplier delay or a burst of cancellations just as customers want their patios ready. If AI agents are going to handle bookings, support or forecasts, a polished demo won’t show how they respond when the pressure is real. Firmulate’s live experiment puts models through a company’s crisis week. Its next step is bringing that exercise to an enterprise’s own business data.
Get pool and patio gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same crises
Firmulate ran frontier AI models through the same small software company and its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped a total: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was not recognizing what to do; it was following through. Firmulate sums up the result as “Same diagnosis, same pitch — no signature.”
The detail buried in the files
The deal depended on a competitor weakness found two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It’s a reminder that a useful agent must connect the evidence already in a business’s records with the moment a decision is due.
The experiment also tested pressure and trust. Fake CEO messages escalated across three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work, missed close
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and let discipline slip, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. The standings therefore tell a broader story than who can analyze a crisis: execution and respect for boundaries matter too.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a snapshot of this experiment, with that difference in setup made explicit.
From watching to testing your own business
The live company gives the experiment a public, watchable setting. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record for every workday. The live page is at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For an enterprise, the proposed next step is a pilot using a read-only export of its own business. Teams can test crisis scenarios against their company and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That moves the question from whether a model can talk convincingly about a business to how it handles that business’s own pressures, evidence and rules.

Watching an AI company makes its choices visible; a pilot can put those choices against your own company’s records and crisis scenarios. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
