
Choosing an AI model for business can look like choosing a phone by its spec sheet: familiar names, polished demos, and a leaderboard. But a new experiment from Firmulate asks a more practical question: when an AI is responsible for a company’s worst week, will it follow through? Moonshot’s Kimi K3 finished second in the Crucible, ahead of three Western frontier models.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate put each frontier model in charge of the same small software company, with the same customers, crises and temptations. The decisions were versioned and auditable. The company is a live experiment with synthetic employees and real money mechanics; its public cash countdown and day-to-day activity can be watched at Firmulate.
The July 2026 Crucible league table puts gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The gap between knowing and doing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That distinction matters for companies evaluating AI agents: recognizing the right answer is not the same as completing the work.
The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that buried fact and closed the deal. Firmulate describes its performance as the cleanest discipline in the field, with one deviation.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
enterprise AI model evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.
The result is a reminder that impressive analysis alone does not settle which model will work best in a company. K3 beat three of four Western frontier models here, but this is one specific wargame, and the ranking is a reason to test models against the work they will actually face—not a universal verdict.
There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Firmulate says its quiz uses 242 real, unedited management decisions, letting visitors guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems. Results and plain-language findings are available at Firmulate’s benchmarks page.

As an affiliate, we earn on qualifying purchases.
Test the work, not just the demo
K3’s second-place finish makes the frontier-model league look open, while the signed-deal gap shows why chat quality can miss what matters in business. If an AI agent may touch your CRM, customer support or forecasts, choosing without testing it on realistic company work is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI model performance benchmarking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
