
The difference between a clever AI and a useful one
Technology buyers are accustomed to watching AI systems summarize documents, draft polished emails and answer questions on command. Firmulate’s live business wargame tests something more consequential: whether an AI agent will investigate the evidence, resist pressure and complete the work needed to produce a commercial result.
In the Crucible League, the decisive clue was not contained in a customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that found it won a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Those that stopped reading lost the deal automatically.
That makes “reads your files before answering” more than a product claim. In this experiment, it became a measurable behavior with a direct business outcome.
As an affiliate, we earn on qualifying purchases.
Same company, same crisis, different outcome
Firmulate runs frontier models as complete companies rather than judging them through isolated chat prompts. Each participant managed the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
The simulated company is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue. Its cash countdown is public, its models have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, live and watchable through Firmulate.
The broad competence of the contestants was not in doubt. Every model identified every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 contract that their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
The missing action matters because the buried fact was available. It simply required the agent to follow the company’s references far enough before responding. Models that did so could support the full-price offer. Models that did not were unable to close, even though they understood the wider situation.
A stress test for trust as well as diligence
The week also included fake CEO messages that escalated over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a crisp explanation for its decision: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean sweep is important. The agents were not merely rewarded for getting things done; they also had to preserve trust while operating under pressure. The do-nothing baseline scored 26 because partial progress still counted, but a single breach of trust capped the total. Firmulate’s standard is blunt: “no amount of good work outweighs a breach of trust.”
The league table rewards follow-through
The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. The full Firmulate benchmark results make the comparison public.
- gpt-5.6-sol found the buried fact and closed the deal.
- Kimi K3 also converted its analysis into a signed agreement.
- Sonnet 5, Fable 5 and Opus 4.8 fell short of the same commercial outcome.
There is an important fairness note: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase its result, but it is useful context for anyone comparing the standings.
Opus 4.8 produced the most striking cautionary tale. It was the most thorough participant, added 80 learned rules and generated the deepest analyses, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. Firmulate observed a weaker version of that discipline problem in the other four participants as well.

As an affiliate, we earn on qualifying purchases.
Buyers should test behavior, not presentation
The lesson for businesses considering AI agents is not that analysis has become unimportant. It is that analysis alone does not complete a workflow. An agent can identify the problem, prepare the pitch, withstand social engineering and still fail at the action that creates value.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz, giving readers another way to examine how difficult attribution can be when outputs are viewed without their outcomes.
For enterprises, the more practical extension is a pilot using a read-only export of their own business. Nothing writes back to real systems. That allows a company to test whether an agent reads the relevant material, respects operational boundaries and follows through before giving it access to a CRM, support queue or forecast.
For gadget and technology enthusiasts, this is a useful reframing of AI performance. The model with the most elaborate answer is not necessarily the model that finishes the job. In Firmulate’s worst-week test, diligent reading was worth a signed €55,000 deal—and failing to read far enough made every other capability irrelevant to that outcome.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.