firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A brilliant chatbot is not necessarily a capable manager

Technology buyers have learned to compare AI models through coding benchmarks and chat arenas. Those tests reveal useful differences in answer quality, but they say little about what happens when an agent faces an overflowing workload, incomplete information and consequences that accumulate across days.

That gap matters if the next generation of agents will touch customer records, support queues or financial forecasts. A polished response is not the same as sound triage. Detecting a problem is not the same as resolving it. And producing an impressive analysis is not the same as telling the board an uncomfortable truth.

Firmulate is testing this distinction through a live, watchable company experiment. Frontier models were asked to run the same small software business through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The result is a case for treating management quality—not chat quality—as a distinct category of AI performance.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between seeing and doing

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a hard principle on trust: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The headline finding was more revealing than the ranking. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the problem neatly: “Same diagnosis, same pitch — no signature.”

This is the measurement gap in miniature. Conventional benchmarks reward the ability to identify the right answer. Business operations reward the ability to carry that answer through to a responsible conclusion. An agent can reason correctly, draft persuasively and still fail the company by leaving the decisive action unfinished.

Reading the company before reacting to the event

The winning clue was not sitting in the incoming customer event. A decisive weakness in the competitor’s position was buried two document references deep inside the company’s own files. Models that followed those references won the deal at full price, worth +€4,583 MRR.

That finding should catch the attention of anyone deploying agents around enterprise data. The important capability was not simply summarisation or recall. It was the managerial habit of checking internal context before acting. In a real organisation, the difference between an adequate response and a valuable one may be hidden in an old account note, a prior decision or a document that is not attached to the latest message.

Trust held up better than execution

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is encouraging. Under explicit social pressure, the field protected approval boundaries and resisted manipulation. It also complicates the familiar fear that capable agents will inevitably sacrifice rules for results. In this experiment, honesty was not the primary weakness. Follow-through was.

Opus 4.8 makes that contrast especially sharp. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, while its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, less strongly, in all four other participants.

That profile challenges the assumption that more visible reasoning necessarily produces better operational performance. Thoroughness can be valuable, but management also requires prioritisation, respect for boundaries and a reliable transition from analysis to action.

A company-shaped curriculum

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. This makes the experiment feel less like a prompt contest and more like an unfolding organisational stress test.

The scenario names point toward a new curriculum for business agents: churn wave, price increase, downround and PR crisis. Each tests judgment across time, competing priorities and institutional trust. Readers can inspect the broader benchmark findings, while 242 real, unedited management decisions also power a public “guess the model” quiz.

There is one important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs beside the result, particularly when the scores are close.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI trust and risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need a management benchmark

The next serious evaluation question is not whether an AI can produce a credible memo. It is whether the agent reads the relevant files, protects trust, escalates when blocked and finishes the work it has justified.

Firmulate is also offering enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That approach turns abstract model comparison into a rehearsal for the conditions an organisation actually cares about.

Coding ability and conversational quality remain useful signals. They are simply incomplete. Before an AI workforce is trusted with customers, cash or forecasts, it should face the corporate equivalent of a fire drill. The strongest agent will not merely explain what management should do. It will do the permitted work, close the loop and remain honest when pressure arrives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business process automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bonsai: Janestreet’s UI Library

Janestreet has released Bonsai, an open-source UI library aimed at improving interface development for financial and data-intensive applications.

Go 1.27 Interactive Tour

Go 1.27 introduces an interactive tour feature aimed at improving developer onboarding and code understanding, announced officially today.

The Future Of AI According To ByteDance’s Founder: No To Distillation

ByteDance’s founder has reportedly banned the use of model distillation in AI development, a move that could impact its future AI strategies and efficiency.

Evaluating AI Performance: DeepSeek-V4-Flash-High’s Ninth Point At A Low Cost

DeepSeek-V4-Flash-High ranks ninth on the Arena leaderboard, with significant performance gains from post-training improvements at a fraction of the cost.