firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Imagine an AI that doesn’t just answer your questions but actually digs through your company’s files to find the critical details that tip the scales of a deal. That’s exactly what a recent live experiment reveals, highlighting a decisive factor that could define the future of AI-powered business decisions.

The Hidden Depths of AI Decision-Making

In a groundbreaking live test, four leading AI models faced the same challenging scenario: guiding a small software company through its worst week, with all the crises, customer demands, and manipulations that would test any manager’s integrity. The goal was simple but revealing: which AI could navigate the storm, stay honest, and close a €55,000 deal?

The results? All four models identified every crisis and refused every attempt at manipulation — a promising sign of their basic ethical and decision-making abilities. But only two of them actually signed the deal based on their analysis. Why? The secret lay buried two document references deep within the company’s files, not in the readily visible customer interactions.

Amazon

enterprise document management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Power of Reading Deep

This experiment’s key insight is subtle but profound: AI agents that delve into a company’s internal documents can uncover the critical details that determine success or failure in negotiations. The models that ignored these deep references missed the full picture, leaving a significant potential deal on the table. Meanwhile, the models that read deeply secured their victory, closing the full €55,000 deal and gaining an additional €4,583 in monthly recurring revenue (MRR).

Amazon

AI document analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond Surface-Level Answers

Many AI demonstrations focus on conversational prowess, but this experiment shows that the real competitive edge lies in context-awareness: does the AI read your files before answering? The models that did so demonstrated a measurable, purchase-deciding property. Without this capability, an AI might look convincing but ultimately miss the critical details that clinch a deal.

Amazon

secure internal file reader for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Under Pressure

In addition to decision-making, the models faced a social engineering test — fake CEO messages escalating over three stages, plus a reporter trick asking for a quick yes/no response. All five models refused to be manipulated, with Kimi K3 explicitly treating the request as a suspected impersonation or approval bypass. This demonstrates that ethical safeguards are not just add-ons but integral to effective AI in business environments.

Amazon

business decision support AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: Reality in Action

The live experiment isn’t just theoretical. Firmulate’s simulated company involves 13 synthetic employees, real money mechanics burning €105k/month against €2.3k MRR, and over 680 self-learned rules. Every decision is versioned and transparent, offering a watchable window into how AI manages real-world business pressures. The site, firmulate.com/live, provides ongoing updates where you can see AI models in action, making strategic decisions that matter.

Insights for Business Buyers

While many might focus on how well an AI writes or chats, the heart of this experiment is about reliability and trustworthiness. Will your AI finish what it starts? Will it read your files thoroughly? Will it stay honest under pressure? These are the questions that matter if AI is to support decisions involving real money and reputation.

The Leaders and Their Scores

  • gpt-5.6-sol scored the highest at 95, successfully uncovering the buried fact and closing the deal, demonstrating complete performance.
  • Kimi K3, the newcomer, scored 93 and was the most disciplined, signing the deal cleanly without slipping.
  • Sonnet 5 scored 88, with a few process slips but still closing the deal.
  • Fable 5 scored 77, also closing but leaving some opportunities on the table due to discipline lapses.

These scores show that even the best models can leave value behind if they don’t read deeply or stay disciplined. The difference is often a matter of whether they look into the company’s internal files before making decisions.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Live Updates: 5 Killed And 5 Injured After Plane With Amazon Logo Crashes At Miami Airport

A plane bearing the Amazon logo crashed at Miami International Airport, killing 5 and injuring 5 others. Investigation ongoing, details still emerging.

Budgeting for Variable Expenses (Bills That Change Monthly)

Wondering how to handle fluctuating bills? Learn effective strategies to budget for variable expenses and stay financially prepared.

30‑Minute Weekly Money Meeting Agenda

Optimize your financial health with a quick 30-minute weekly money meeting agenda that reveals essential tips to stay on track.

A Smart Budget for Home Fitness Without Buying Junk

Keen on building an effective home gym without wasteful purchases? Discover smart tips to optimize your fitness budget and stay motivated.