
In a world increasingly driven by artificial intelligence, the question is often reduced to: can machines write convincingly? But a recent live experiment challenges that notion, revealing that volume and diligence alone don’t guarantee success—especially when trust and prioritization are on the line.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Real Test: Beyond Just Spotting Crises
At the heart of a groundbreaking experiment conducted by the public AI benchmarking platform Firmulate, four advanced AI models were tasked with running a simulated small software company through its worst week. The goal was straightforward yet revealing: see if these models could identify crises, navigate temptations, and, most importantly, close a critical deal worth €55,000.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Crises, Different Results
Each AI model experienced identical scenarios—same customers, same crises, same attempts at manipulation. The models’ decisions were meticulously versioned and auditable, ensuring transparency in every move. The results? All four models flagged every crisis and refused every manipulation attempt. Yet, only two managed to secure the deal, despite all four diagnosing the same issues and presenting the same pitch.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deeper Files
The decisive factor was not the surface-level crisis detection but the AI’s ability to read deeper into the company’s own documentation. The winning models, Kimi K3 and GPT-5.6, uncovered a critical piece of information buried two document references deep—information that made the difference between closing the deal and walking away empty-handed. The models that read the full file secured the deal at full price, adding approximately €4,583 monthly recurring revenue (MRR).
As an affiliate, we earn on qualifying purchases.
Discipline and Focus Matter
Firmulate’s experiment also examined how disciplined each AI was in handling requests. For instance, when presented with social engineering attempts—fake CEO messages escalating over three stages plus a reporter trick—all models refused. Kimi K3 justified its refusal by treating the requests as potential impersonation or approval bypass, demonstrating disciplined skepticism.
As an affiliate, we earn on qualifying purchases.
The Human-Like Flaws: Overload and Slips
Despite their sophistication, the most thorough participant, Opus 4.8, with over 80 learned rules and deepest analyses, finished last in the final tally. Its downfall was a slip in discipline—deciding to write attempts into a restricted department rather than escalating issues appropriately. This illustrates a crucial insight: diligence alone doesn’t guarantee impact. Overloading on rules and volume can lead to missed priorities, as discipline and strategic focus matter more.
Implications for Business and AI Deployment
What does this mean for enterprises considering AI solutions? The experiment underscores that AI’s ability to produce convincing language or handle surface tasks isn’t enough. Success hinges on whether AI can see the full picture—reading relevant files, maintaining discipline under pressure, and making prioritized decisions. The models that went deeper into documents and maintained disciplined refusal outperformed their peers in closing deals that mattered.
The Fairness Note: Different Settings, Same Weakness
The Kimi K3 model ran without an effort parameter (default API setting), while others ran at high effort levels. Yet, all models displayed similar weaknesses, reinforcing the idea that raw diligence isn’t enough—focused prioritization and trustworthiness are key.
What This Means for Your Business
Before deploying AI into critical decision-making roles—whether in customer support, CRM, or forecasting—businesses should ask: does the AI finish what it starts? Does it read your files comprehensively? Will it stay honest under pressure? The experiment demonstrates that the true value of AI lies not in how well it writes or how much it processes, but in whether it can deliver impactful, trustworthy work.
See the Live Experiment in Action
Firmulate’s platform offers a transparent, real-time view of these AI “companies” running through crises, making decisions that mimic real business pressures. You can watch the live scenarios unfold at firmulate.com/live and explore the full benchmark results at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.