
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
How AI Is Testing Its Mettle in the Business World
Imagine a real company running every workday, with AI models steering decisions through crises, temptations, and strategic choices—all in front of your eyes. Now, what if some models perform better than others, even under pressure? That’s exactly what a groundbreaking experiment by Firmulate has revealed about the state of AI management today.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI Models to the Test
In July 2026, four prominent AI frontier models were challenged to run the operations of a small software company during its worst week. Each model faced identical crises, customer interactions, and temptation scenarios, with decisions thoroughly tracked and auditable. This wasn’t a demo—this was a real-time, live experiment designed to measure management quality, not just chat prowess.
The Models and Their Scores
- gpt-5.6-sol scored the highest with 95 points, successfully identifying a buried fact in the company’s files and sealing a €55,000 deal, generating an additional €4,583 in monthly recurring revenue (MRR).
- Kimi K3, the newcomer from Moonshot, finished just behind at 93 points. It demonstrated the cleanest discipline, resisted manipulations, and secured the same deal.
- Sonnet 5 scored 88, also closing the deal but with minor slips in process discipline.
- Fable 5 and Opus 4.8 scored 77 and 73 respectively, with Opus showing the weakest discipline and leaving the close on the table.
Significantly, all models identified crises and refused manipulative social engineering attempts, including fake CEO messages and reporter tricks. The decisive factor in winning the deal was reading a document reference deep in the company’s internal files—a subtle but critical insight.
As an affiliate, we earn on qualifying purchases.
Why the Leader Matters: Beyond Chat Quality
This live experiment underscores an essential truth for anyone relying on AI in business: success isn’t just about how well an AI can generate language in demos. It’s about whether it can finish what it starts, stay honest under pressure, and leverage internal data effectively. The leader, gpt-5.6-sol, demonstrated the ability to find buried information and close deals—an indication of management competence that surpasses simple chat interactions.
Discipline and Trust in AI Decisions
Among the tested models, only Kimi K3 ran without an effort parameter—meaning it operated at the API default setting—yet still managed to outperform many peers. Meanwhile, Opus 4.8, despite its deep rule learning and thorough analysis, failed to close the deal, highlighting that thoroughness alone doesn’t guarantee execution quality.
AI business decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How This Changes the AI Landscape
Traditional benchmarks often focus on chat quality or language fluency, but this experiment from Firmulate reveals that the real measure of an AI’s usefulness in management is its integrity, decisiveness, and ability to read and interpret internal data. As these models are poised to influence CRM, customer support, and forecasting, choosing one that can consistently finish its work under pressure is crucial.
As an affiliate, we earn on qualifying purchases.
The Open League and the Future of AI Management
With scores ranging from 95 to 73, the league shows a clear competitive landscape: the top models can identify critical information and close deals, whereas even the most thorough competitors can falter on execution. This open league suggests that picking an AI model without your own rigorous testing is a gamble.
Watch the Experiment Live and Learn
The entire process is observable at firmulate.com/live, where you can see how real companies, with real money mechanics and self-learned rules, are managed by AI—every decision, every crisis, every outcome. This isn’t just a test; it’s a glimpse into the future of AI-driven management.

Key Takeaway
In a real business environment, the AI model’s ability to find hidden insights, stay disciplined, and finish tasks under pressure matters more than just chat quality. The recent experiment shows that a newcomer can outperform established models when tested rigorously, emphasizing the importance of real-world validation before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
