AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A lesson in reading beyond the obvious

In science and education, good judgment often starts with a simple habit: look for the evidence others missed. Firmulate put that habit to a business test. Several AI models faced the same crisis week at the same small software company, and the pivotal clue was buried two document references deep in the company’s own files.

The experiment is real and watchable at Firmulate. Its results offer a practical question for organizations considering AI: can a model carry its analysis through to a sound decision when the stakes are real?

One company, one difficult week

Firmulate’s Crucible League gave each frontier model the same small software company, customers, crises and temptations. The experiment tracked decisions as they happened, making the work auditable. The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26.

The broad result was striking: every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The company’s records held a decisive competitor weakness, but it sat two document references away from the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.

That gap between seeing and doing is captured in the experiment’s summary: “Same diagnosis, same pitch — no signature.” A model can identify the situation and make a persuasive case, yet still leave the consequential step unfinished.

Trust under pressure

The models also faced staged social engineering: fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning described the request as a “suspected approval-bypass / possible impersonation.” The do-nothing baseline’s 26 points reflected partial progress, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The participant profiles showed other differences. Opus 4.8 was the most thorough, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and attempted to write into a locked department instead of escalating. That weakness appeared, less strongly, in all four models. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh—context readers should keep in mind when comparing the results.

From watching to a company-specific pilot

Firmulate’s live company gives the benchmark a visible setting: 13 synthetic employees operate with real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. The company burns €105k a month against €2.3k MRR, and every workday is versioned. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

For an enterprise, the next step is a pilot against its own business. Firmulate describes using a read-only export to stage crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. The pilot does not write back to real systems. That creates a way to examine how AI might handle a company’s customers, pipeline and rules before entrusting it with live operations.

The experiment’s lesson is concrete: spotting a crisis and respecting boundaries are essential, but organizations also need to know whether a model can find relevant evidence and complete the work its own analysis supports. A company-specific wargame can put those questions in context.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s results show why AI readiness is more than a contest in polished answers. Models can recognize trouble and resist manipulation, yet still differ in whether they uncover a buried clue, close a justified deal and follow the right escalation path. A pilot lets an enterprise explore those behaviors against its own scenarios using a read-only export, with nothing writing back to real systems.

To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Monitoring Soil Carbon With MRV Technologies: Opportunities and Challenges

Learning about MRV technologies for soil carbon reveals promising opportunities and challenges that are essential for effective management.

Climate Forecasting Tools for Farmers: Harnessing AI and Satellite Data

Absolutely! Discover how AI-driven satellite climate forecasting tools can revolutionize your farming strategies and help you stay ahead of weather challenges.

Soil Health Monitoring: Sensors and Data Analytics

Unlock the potential of soil health monitoring with sensors and analytics to optimize your farming practices and ensure sustainable land management.