firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model for business can look like choosing a phone by its spec sheet: familiar names, polished demos, and a leaderboard. But a new experiment from Firmulate asks a more practical question: when an AI is responsible for a company’s worst week, will it follow through? Moonshot’s Kimi K3 finished second in the Crucible, ahead of three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate put each frontier model in charge of the same small software company, with the same customers, crises and temptations. The decisions were versioned and auditable. The company is a live experiment with synthetic employees and real money mechanics; its public cash countdown and day-to-day activity can be watched at Firmulate.

The July 2026 Crucible league table puts gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That distinction matters for companies evaluating AI agents: recognizing the right answer is not the same as completing the work.

The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that buried fact and closed the deal. Firmulate describes its performance as the cleanest discipline in the field, with one deviation.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

enterprise AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.

The result is a reminder that impressive analysis alone does not settle which model will work best in a company. K3 beat three of four Western frontier models here, but this is one specific wargame, and the ranking is a reason to test models against the work they will actually face—not a universal verdict.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Firmulate says its quiz uses 242 real, unedited management decisions, letting visitors guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems. Results and plain-language findings are available at Firmulate’s benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not just the demo

K3’s second-place finish makes the frontier-model league look open, while the signed-deal gap shows why chat quality can miss what matters in business. If an AI agent may touch your CRM, customer support or forecasts, choosing without testing it on realistic company work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Devtools Must Be Open Source

Industry advocates argue that developer tools should be open source to promote transparency, security, and community collaboration.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX acquired AI coding tool Cursor for $60 billion in stock, a move seen as a bargain due to rapid revenue growth and strategic advantages.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a full publishing workflow that stays on your machine. No cloud dependency, total control, and privacy-focused.

Fable and Mythos: How Anthropic Shipped Its Most Powerful Model to Everyone

Anthropic launches Fable 5, a highly capable AI model with safety features allowing broad access, while its Mythos 5 version remains restricted for security reasons.