
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
AI’s Real-World Test Reveals Surprising Gaps in Trust and Performance
In an unprecedented live experiment, leading AI models were tasked with running a small software company through its toughest week. The results? While all recognized crises and resisted manipulative tactics, only two managed to close a critical deal, highlighting not just technical prowess but also the importance of integrity and thoroughness in AI decision-making.
The Live Wargame: Putting AI to the Test in a Simulated Business Environment
Developed by Firmulate, this experiment involves running AI models as complete, decision-making companies facing real crises, customer temptations, and internal challenges. Every decision is recorded and auditable, ensuring transparency in performance. The models faced the same scenario: a week of crises including customer churn risks, security challenges, and social engineering attempts designed to test their integrity.
The models tested included the latest frontier AI solutions, such as gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8. Each was evaluated on their ability to spot crises, refuse manipulative offers, and ultimately, to close a vital €55,000 deal based solely on their analysis, without external influence or shortcuts.
Clear Results: Performance, Integrity, and Business Outcomes
The scores out of 100 reveal meaningful distinctions: gpt-5.6-sol scored the highest at 95, followed closely by Kimi K3 at 93, then Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Notably, the do-nothing baseline scored 26, underscoring how far these models have advanced in handling complex, realistic business scenarios.
All models successfully identified every crisis and resisted manipulation attempts, including social engineering scenarios where fake CEO messages escalated over three stages. For example, all models refused to approve suspicious requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Hidden Weakness: Reading Deep into Files Matters
The decisive factor in who won the deal was whether the model read into the company’s internal files. The winner, gpt-5.6-sol, and the runner-up, Kimi K3, both read two document references deep into the files to uncover a buried security issue. This allowed them to win the deal at full price, adding €4,583 MRR to the fictional company’s revenue.
In contrast, the other models, including Opus 4.8, missed this critical piece of information, leaving the deal on the table despite their analytical depth. This showcases that technical skill alone isn’t enough; thoroughness and attention to internal data are crucial in real-world AI decision-making.
Discipline Under Pressure and the Cost of Missed Opportunities
The experiment also highlighted discipline differences. Opus 4.8, the most thorough participant with over 80 rules learned and deep analysis, finished last in closing the deal. Its weakness was in overlooking critical internal documents, instead writing attempts into a locked department—indicative of a lack of escalation discipline. Interestingly, the same gap appeared, albeit weaker, in the other models, suggesting a common challenge in balancing thoroughness and process discipline.
Fairness and Transparency in Evaluation
It’s important to note that Kimi K3 ran without an effort parameter (the API default), while the other models operated at xhigh effort, ensuring fair comparison. The experiment’s transparency—full decision versioning and auditable decision paths—underscores the importance of trustworthiness in AI performance.
Implications for Business and AI Adoption
This live experiment by Firmulate demonstrates that AI models are increasingly capable of managing complex, high-stakes business scenarios. The key takeaway? Effectiveness isn’t just about generating responses but about making honest, thorough, and strategic decisions. For enterprises considering AI solutions, demonstrating the ability to finish what it starts—reading relevant internal data and resisting manipulations—will determine success and ROI.
Watching this experiment unfold live at firmulate.com/live offers a real glimpse into the future of AI in management. It’s not merely a matter of chat quality but of trustworthy, disciplined work that can impact your bottom line.

AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaway: Trust, thoroughness, and discipline matter more than ever in AI-driven business management. The ability to read deeply, resist manipulation, and close deals under pressure sets the top models apart.
As the league table shows, the best AI models are closing the gap with human-like decision-making, but the real challenge is ensuring integrity and thoroughness in every decision. For business leaders, the message is clear: pick your AI partner carefully—one that can finish what it starts, read your files deeply, and stay honest under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI data analysis tools for internal files
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and fraud detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI deal-closing automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
