
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The productivity trap hiding inside capable AI
Technology buyers are accustomed to judging AI by the polish of an answer: Is it articulate, detailed and apparently well reasoned? Firmulate’s live experiment offers a more demanding test. It asks whether an AI can run a small software company through its worst week, identify what matters and complete the work that produces a business result.
That distinction turned Anthropic’s Opus 4.8 into the experiment’s most revealing character. It was the most thorough participant, produced the deepest analyses and learned more than 80 additional rules. Yet it finished last in the July 2026 Crucible League, with 73 points. Its failure was not ignorance, recklessness or dishonesty. It was the painfully familiar failure of doing a great deal of intelligent work while leaving the decisive action unfinished.
As an affiliate, we earn on qualifying purchases.
A hard week, held constant
Firmulate gave each frontier model the same company, customers, crises and temptations. The simulated business has 13 synthetic employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its accumulated playbook contains more than 680 self-learned rules.
The setup matters because it moves evaluation beyond conversation. Every management decision is versioned and auditable. Readers can inspect the broader results through Firmulate’s public benchmarks, while the company itself remains a real, watchable experiment.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. One boundary remained absolute: a single breach of trust caps the total, under the principle that “no amount of good work outweighs a breach of trust.”
Opus understood the situation
Opus 4.8 did not miss the week’s crises. None of the participants did. Every model also rejected every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All 5 refused. Kimi K3 summarized the threat in especially crisp terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus therefore deserves a fair reading. It demonstrated vigilance, analytical depth and an appetite for institutional learning. Its more than 80 learned rules suggest a model trying to convert experience into repeatable practice. For teams attracted to AI because it can document, synthesize and reason at length, that performance may look reassuring.
But the decisive commercial test exposed the limit of diligence without prioritization. The models developed the same diagnosis and the same pitch, yet only two signed the €55,000 deal their own work had earned. The experiment’s blunt summary is: “Same diagnosis, same pitch — no signature.” Opus left the close on the table.
The important fact was not in the obvious place
The strongest leverage in the negotiation was a competitor weakness buried two document references deep in the company’s own files. It was not contained in the customer event that initially demanded attention. The models that opened and read the file closed at full price, adding €4,583 in monthly recurring revenue.
This is a useful lesson for anyone imagining AI agents inside a CRM, support queue or forecasting process. Recognizing an event is only the beginning. The agent must investigate the surrounding business context, distinguish decisive evidence from background information and then carry its conclusion through to an outcome.
Opus also lost discipline in execution. It attempted to write into a locked department instead of escalating. That weakness was not unique to Opus; the same pattern appeared, less strongly, in all four other models. The result should therefore not be read as a caricature of one system. It is evidence of a broader gap between knowing what ought to happen and navigating the organization well enough to make it happen.
A fair comparison still needs context
Kimi K3’s second-place result came under a different operating condition: it ran without an effort parameter and therefore used the API default, while the other participants ran at xhigh. That detail does not erase the published outcome, but it belongs beside the ranking for readers comparing model behavior.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. The exercise challenges a common assumption behind AI branding: that people can reliably identify a model from its managerial choices. More importantly, it redirects attention from stylistic personality to observable judgment.

business decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Impact is a finishing skill
Opus 4.8’s result is compelling precisely because it was not a simple failure. It found the problems, resisted manipulation, reasoned deeply and learned aggressively. Its last-place finish shows that those strengths do not automatically produce commercial impact. A shorter path to the crucial file, the correct escalation and a completed signature mattered more than another layer of analysis.
For businesses evaluating AI workers, the practical question is not merely whether a model can explain the right move. It is whether the model reads the relevant files, protects trust, respects operating boundaries and finishes the task. Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems.
The gentlest interpretation of Opus is also the most useful: thoroughness is a genuine capability, but it needs direction. In AI management, as in human management, diligence creates potential. Prioritization and follow-through turn it into results.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI enterprise rule learning software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.