
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What happens when a workplace is under pressure?
In health and wellness, people know that good intentions can falter when the day gets hard. The same question applies to AI agents being considered for consequential business work: can they spot trouble, protect trust and follow through when several demands collide? Firmulate is testing that question in a live, watchable experiment—and now invites businesses to try the exercise with their own data.
A company’s worst week, repeated
For its final Crucible League in July 2026, Firmulate gave frontier models the same small software company to run through its worst week. Customers, crises and temptations were held constant; only the model changed. Decisions were versioned and auditable, so observers can follow what each participant chose.
All five models in the final league did better than the do-nothing baseline, which scored 26. The leaderboard ended with gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The benchmark also makes integrity consequential: one breach of trust caps the total. Its principle is plain: “no amount of good work outweighs a breach of trust.”
Recognizing a crisis is only part of the job
The striking finding was not that models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” A model may give a persuasive account of what should happen and still leave the opportunity untouched.
The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a practical reminder for any organization: the answer can depend on whether an agent consults the information already available, then acts on what it finds.
Integrity faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” These are concrete signals of caution under pressure, though the deal results show that caution and completion must be assessed together.
More analysis did not guarantee a better result
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted writes in a locked department instead of escalating. The same weakness appeared, less strongly, in all four. Firmulate also notes a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
The public live company gives the experiment a continuing setting. It has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Firmulate says the site rebuilds itself twice a day. Readers can also explore a “guess the model” quiz built from 242 real, unedited management decisions.
From watching to a company-specific pilot
A general benchmark can show how models behave in one shared scenario. A pilot asks a closer-to-home question: how would they handle your business’s customers, operating rules and crisis scenarios? Firmulate says enterprises can run the wargame against a read-only export of their business and receive a board report with model rankings and weak points in their playbooks. The exercise does not write back to real systems.
That boundary matters when testing agents intended for work involving customer records, support queues or forecasts. A read-only trial lets a company examine decisions against its own context before considering deployment. The findings here suggest what to look for: not just whether a model identifies a problem or refuses a manipulation, but whether it consults relevant records, completes justified work and escalates when it cannot proceed.

Test judgment before handing over responsibility
Firmulate’s experiment shows a gap between seeing the right move and carrying it through. For leaders considering AI agents, a wargame against company-specific scenarios offers a way to inspect that gap before an agent touches live systems. Explore the live experiment at firmulate.com. To discuss a read-only pilot for your business, visit the pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
