firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

How AI Is Testing Its Mettle in the Business World

Imagine a real company running every workday, with AI models steering decisions through crises, temptations, and strategic choices—all in front of your eyes. Now, what if some models perform better than others, even under pressure? That’s exactly what a groundbreaking experiment by Firmulate has revealed about the state of AI management today.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: Putting AI Models to the Test

In July 2026, four prominent AI frontier models were challenged to run the operations of a small software company during its worst week. Each model faced identical crises, customer interactions, and temptation scenarios, with decisions thoroughly tracked and auditable. This wasn’t a demo—this was a real-time, live experiment designed to measure management quality, not just chat prowess.

The Models and Their Scores

  • gpt-5.6-sol scored the highest with 95 points, successfully identifying a buried fact in the company’s files and sealing a €55,000 deal, generating an additional €4,583 in monthly recurring revenue (MRR).
  • Kimi K3, the newcomer from Moonshot, finished just behind at 93 points. It demonstrated the cleanest discipline, resisted manipulations, and secured the same deal.
  • Sonnet 5 scored 88, also closing the deal but with minor slips in process discipline.
  • Fable 5 and Opus 4.8 scored 77 and 73 respectively, with Opus showing the weakest discipline and leaving the close on the table.

Significantly, all models identified crises and refused manipulative social engineering attempts, including fake CEO messages and reporter tricks. The decisive factor in winning the deal was reading a document reference deep in the company’s internal files—a subtle but critical insight.

Amazon

AI data analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Leader Matters: Beyond Chat Quality

This live experiment underscores an essential truth for anyone relying on AI in business: success isn’t just about how well an AI can generate language in demos. It’s about whether it can finish what it starts, stay honest under pressure, and leverage internal data effectively. The leader, gpt-5.6-sol, demonstrated the ability to find buried information and close deals—an indication of management competence that surpasses simple chat interactions.

Discipline and Trust in AI Decisions

Among the tested models, only Kimi K3 ran without an effort parameter—meaning it operated at the API default setting—yet still managed to outperform many peers. Meanwhile, Opus 4.8, despite its deep rule learning and thorough analysis, failed to close the deal, highlighting that thoroughness alone doesn’t guarantee execution quality.

Amazon

AI business decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How This Changes the AI Landscape

Traditional benchmarks often focus on chat quality or language fluency, but this experiment from Firmulate reveals that the real measure of an AI’s usefulness in management is its integrity, decisiveness, and ability to read and interpret internal data. As these models are poised to influence CRM, customer support, and forecasting, choosing one that can consistently finish its work under pressure is crucial.

Amazon

AI enterprise management models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Open League and the Future of AI Management

With scores ranging from 95 to 73, the league shows a clear competitive landscape: the top models can identify critical information and close deals, whereas even the most thorough competitors can falter on execution. This open league suggests that picking an AI model without your own rigorous testing is a gamble.

Watch the Experiment Live and Learn

The entire process is observable at firmulate.com/live, where you can see how real companies, with real money mechanics and self-learned rules, are managed by AI—every decision, every crisis, every outcome. This isn’t just a test; it’s a glimpse into the future of AI-driven management.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaway

In a real business environment, the AI model’s ability to find hidden insights, stay disciplined, and finish tasks under pressure matters more than just chat quality. The recent experiment shows that a newcomer can outperform established models when tested rigorously, emphasizing the importance of real-world validation before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

CG Oncology Announces Publication Of Pivotal Phase 3 BOND-003 Cohort C Study Results In The Lancet Oncology

CG Oncology announces publication of pivotal Phase 3 BOND-003 Cohort C study results in The Lancet Oncology, highlighting key findings on treatment efficacy.

Energiewende: Anteil Erneuerbarer Energien An Deutscher Stromerzeugung Auf Rekordhoch

Germany’s renewable energy share in electricity production has reached an all-time high, marking a significant milestone in the Energiewende transition.

National Airlines Receives IATA IEnvA Certification At IATA Headquarters

National Airlines has received IATA’s IEnvA certification, marking a significant step in sustainable aviation practices. Certification was awarded at IATA headquarters.

What a Full Standing Desk Setup Really Costs

Beyond the initial price, understanding the true costs of a full standing desk setup can help you make smarter, more comfortable choices—so let’s explore further.