firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine running a business where your AI workforce faces the same crises, temptations, and tough choices as a human team. Would it succeed? Would it cheat? Would it make smarter decisions? At Firmulate, real frontier AI models are tested in a live environment, revealing fascinating differences in their management styles and integrity. This isn’t just theory — it’s an ongoing experiment with tangible results that could reshape how we think about AI in business.

The Live Business Wargame: Putting AI to the Test

At the heart of the experiment is a small software company turned into a testing ground for AI management. Four leading frontier models — GPT-5.6, Kimi K3, Sonnet 5, and Fable 5 — each navigate the same challenging week. From demanding customers and ethical dilemmas to manipulative ploys, every crisis is identical across the board, ensuring a fair comparison.

All models demonstrated a remarkable ability: they spotted every crisis and refused every attempt at manipulation. For example, when fake CEO messages were escalated in stages, all five models refused to engage, citing concerns over impersonation and fraud. Their integrity was clear: honesty under pressure is a core trait.

Amazon

AI management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Decision-Making and Outcomes

The results? Only two models, GPT-5.6 and Kimi K3, signed a €55,000 deal that their own analyses had earned — representing real, valuable revenue. Despite identical diagnoses and pitches, the other two models, Sonnet 5 and Fable 5, left money on the table, declining to finalize the deal. Interestingly, the decisive factor wasn’t the immediate crisis but something buried two document references deep in the company’s files, which the successful models read and understood. Access to and comprehension of these internal documents gave GPT-5.6 and Kimi K3 a clear edge, allowing them to close at full price (+€4,583 MRR).

Amazon

business AI decision making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Personality of AI Managers

Beyond raw performance, the models exhibit measurable management personalities. GPT-5.6 is the most thorough, analyzing deeply and leaving no stone unturned. Kimi K3 is disciplined and fair, running without an effort parameter and maintaining strict integrity. Sonnet 5 is more process-oriented, with some slips, while Fable 5 shows a tendency to leave deals unclosed, slipping into departmental silos instead of escalating issues. These differences highlight that, like humans, AI models can develop distinct management styles, influencing their success in complex scenarios.

Amazon

AI ethics and integrity management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Business?

For companies considering AI for management roles, the implications are clear: it’s not just about how well the AI writes or communicates, but whether it can finish what it starts, read relevant documents thoroughly, and resist manipulative pressures. The live experiment underscores that AI models can be trusted to identify crises and refuse unethical shortcuts — but their management personality affects outcomes. The current league table places GPT-5.6 at the top, with a score of 95, followed closely by Kimi K3 at 93, demonstrating their reliability and integrity.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Money, Real Time, Real Decisions

The experiment is visible in real-time at firmulate.com/live. The company runs every weekday with 13 synthetic employees managing real money mechanics — burning €105k/month against €2.3k MRR. Every decision, every crisis, every attempt at manipulation is logged and auditable, providing a transparent view into how these models perform under pressure. It’s a test bed for the future of AI in management, risk assessment, and operational integrity.

Why You Should Care: The Future of AI in Business

If AI agents will eventually engage with your CRM, customer support, or forecasting tools, the key questions aren’t how well they chat, but whether they finish tasks, read critical files, and stay honest under pressure. This experiment shows that current frontier models can do these things well, but their management style matters. Choosing the right model could mean the difference between sealing a lucrative deal or leaving money on the table.

Try It Yourself

Business leaders and decision-makers can run their own wargames against their enterprise data, using the same format of the experiment, without risking their actual systems. Visit firmulate.com/pilot.html to learn how to simulate your company’s worst week and gauge your AI’s managerial integrity and decision-making depth.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Budgeting for Variable Expenses (Bills That Change Monthly)

Wondering how to handle fluctuating bills? Learn effective strategies to budget for variable expenses and stay financially prepared.

50‑Week Savings Challenge Tracker

Keep your savings goals on track with a 50-week challenge tracker that simplifies progress monitoring and motivates you to reach your financial dreams.

What Backyard Delivery and Assembly Costs Really Look Like

Costs for backyard delivery and assembly vary widely depending on several factors, and understanding these details can help you budget effectively.

Budgeting Tools Comparison: Apps Vs Spreadsheets Vs Paper

AIThis post was created with the assistance of artificial intelligence (AI).Choosing the…