firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a manager by giving them a week of the worst possible crises—then seeing if they can navigate, read the situation, and trust their team. Now, picture doing that with AI models. The results might surprise you: even the ‘do-nothing’ approach scores 26 out of 100. What does that tell us about measuring AI—and the importance of trust and thoroughness in business decisions?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real Test: Simulating a Business Crisis

At the heart of the recent Firmulate experiment is a simple idea: treat AI like a manager facing a challenging week at a small software company. The scenario is consistent for every model—same customers, same crises, same temptations to cheat or manipulate. Each AI model is tasked with making decisions that affect real money mechanics, customer trust, and operational discipline.

This approach is transparent and auditable, meaning every choice and decision can be tracked and reviewed. The goal isn’t just to see if AI can produce convincing chat responses but whether it can manage the complex, often messy realities of running a business.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Unexpected Findings: Scores and Trust

One key insight is the baseline score: a ‘do-nothing’ approach—essentially, not making any decisions—receives a score of 26 out of 100. This might seem odd at first, but it illustrates a crucial point: partial progress counts, and even minimal engagement can keep the AI from failing outright. More importantly, a single breach of trust, such as signing a deal without proper analysis, caps the maximum score at that level. In other words, no amount of good work can outweigh a fundamental breach of integrity.

Amazon

business crisis management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Model Performance: Discipline and Diligence

Among the models tested, the highest scorer was GPT-5.6-sol, which scored 95 points, successfully reading the crucial buried fact in company files and closing a real deal. Close behind was Kimi K3 with 93 points, showcasing the best discipline among the field—refusing manipulative asks and sticking to verified facts. Sonnet 5 followed with scores of 88 and 77, demonstrating that even close performance often involved slips or missed opportunities.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Deep into Files

Interestingly, the decisive advantage for the top models was their ability to read two documents deep into the company’s files—an area where many models faltered. Those that successfully identified and used this information won the deal at full price, adding €4,583 MRR to the company’s revenue. This highlights that thoroughness—reading and verifying information—is more critical than surface-level interactions or quick chat responses.

Amazon

trust and discipline management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust in Manipulative Scenarios

Another test was social engineering, where models faced staged requests from a fake CEO escalating over three stages, plus a reporter trick asking for a simple yes/no answer. All models refused to manipulate or sign off on dubious requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline under pressure is essential for real-world AI deployment, especially in sensitive management or customer support roles.

The Live Business Simulation

The experiment wasn’t just theoretical. It ran on a synthetic but realistic business—13 employees, real money mechanics, a €105k/month burn rate, and a public cash countdown. Every day, the AI models made decisions within a complex environment, with over 680 self-learned rules guiding their actions. The entire operation is publicly observable at firmulate.com/live.

Lessons for Business Leaders

This experiment underscores a vital truth: AI models are not just about generating convincing chat or natural language—they’re about making consistent, honest, and effective decisions. A model that reads deeply, refuses manipulation, and stays disciplined can close real deals and manage crises effectively. Conversely, even a thorough model can fail if it slips on discipline or trustworthiness.

For enterprise decision-makers, the takeaway is clear: before hiring an AI workforce, use tools like Firmulate’s live wargames to test how models perform under pressure—just like a real manager. It’s not enough for AI to sound good; it must deliver trustworthy, actionable results.

The Benchmark and Its Significance

The current league table reveals that no model scores below 26 points—the baseline score for doing nothing—highlighting the importance of partial progress and minimal engagement. The best models score in the low 90s, but what matters most is their ability to avoid breaches of trust and to dig into critical information—skills that are often overlooked in standard AI demos.

Ultimately, this transparent, auditable benchmark sets a new standard for evaluating AI in management roles—emphasizing discipline, thoroughness, and integrity over superficial performance. For businesses, it offers a clearer picture of how AI can genuinely support decision-making in high-stakes environments.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

AI management models are only as good as their discipline and thoroughness. The Firmulate benchmark shows even a do-nothing approach scores 26, but top models excel by reading deeply, refusing manipulation, and maintaining trust—crucial qualities for real-world deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Weekend Project Budget That Gets Out of Control Fast

Many weekend project budgets spiral out of control quickly—discover crucial tips to keep expenses in check before it’s too late.

No-Budget Budgeting: A Simple No-Fuss Approach

Part of a stress-free money management plan, no-budget budgeting helps you stay flexible and aware—discover how it can transform your financial habits.

AI‑Powered Budgeting Assistants: Worth It?

Budgeting assistants powered by AI can transform your financial management— but are they truly worth the potential risks and benefits?

Using Multiple Bank Accounts to Organize Your Budget

Boost your budgeting with multiple bank accounts—discover how clear boundaries and organized finances can transform your financial future.