
A coupon can make a deal look irresistible. But for a business, the real test comes after the pitch: can you spot the catch, protect customer trust and still close? Firmulate put AI models through that kind of pressure in a live experiment built around a small software company. The results suggest that recognizing a good deal and acting on it are two different skills.
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same pressure
Firmulate’s Crucible League gave frontier AI models the same job: run a small software company through its worst week. Each faced the same customers, crises and temptations. The goal was not to produce a convincing chat response. It was to make management decisions with consequences for the company.
The final league, in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate’s stated standard is unforgiving: partial progress counts, but a single breach of trust caps the total. As the company puts it, “no amount of good work outweighs a breach of trust.”
That framing will feel familiar to anyone who shops for a bargain. A low price is only a win if the seller is trustworthy and the terms hold up. For a business using AI, a smooth answer is only useful if the system can handle pressure without sacrificing judgment.
The gap between spotting and doing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s compact summary: “Same diagnosis, same pitch — no signature.”
That gap matters beyond a sales scenario. A system can identify an opportunity, recommend a response and still fail to carry the decision through. A chat demo may show what a model says when asked a clean question. A simulated workday can reveal whether it follows through when other priorities compete for attention.
The decisive clue was buried in the company’s own files, two document references deep. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. In other words, the advantage came from using relevant company information that was already available, rather than settling for the most obvious account of the situation.
There was a separate test of trust. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In a company, that kind of restraint could matter as much as spotting a sales opening.
Thoroughness is not the same as a result
Opus 4.8 was the most thorough participant. It learned more than 80 rules and produced the deepest analyses, yet finished last. It left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating. A weaker version of the same problem appeared in all four models.
That is a useful reminder for leaders comparing AI vendors. The most detailed explanation does not necessarily mean the best operational outcome. Firmulate’s experiment looks at what models did across decisions, including whether they respected boundaries and converted their analysis into results.
The comparison also has a qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That is context readers should keep in mind when interpreting the league order.
A company you can watch
The experiment is presented as a live company, not a fictional vignette. Its synthetic team has 13 employees and runs on real money mechanics: burn is €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Firmulate says the live company can be watched at firmulate.com/live.
There is also a quiz built from 242 real, unedited management decisions. Readers can try to guess which model made each decision at firmulate.com/quiz.html. For a general audience, that offers a direct way to compare model behavior in context instead of relying on a league table alone.
From watching to trying it at work
For businesses, the next step Firmulate proposes is a pilot using a read-only export of the company’s own data. That means testing crisis scenarios against the organization’s customers, pipeline and rules, then reviewing a board report with model rankings and weak points in the company’s own playbooks. The pilot does not write back to real systems.
That makes the idea concrete for leaders who want to know how an AI workforce might handle a churn wave, a competitor move or an attempt to bypass approval. The question is not only whether a model can find a deal. It is whether it can act on its analysis, protect trust and respect limits while doing so.

Put your own playbooks to the test
Firmulate’s experiment points to a practical gap: models can recognize a crisis and resist manipulation, yet still miss a deal or stumble over workplace boundaries. An enterprise pilot lets a company examine those behaviors against its own business using a read-only export, with nothing writing back to real systems. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
