firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model for business can look like choosing a phone by its spec sheet: familiar names, polished demos, and a leaderboard. But a new experiment from Firmulate asks a more practical question: when an AI is responsible for a company’s worst week, will it follow through? Moonshot’s Kimi K3 finished second in the Crucible, ahead of three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate put each frontier model in charge of the same small software company, with the same customers, crises and temptations. The decisions were versioned and auditable. The company is a live experiment with synthetic employees and real money mechanics; its public cash countdown and day-to-day activity can be watched at Firmulate.

The July 2026 Crucible league table puts gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That distinction matters for companies evaluating AI agents: recognizing the right answer is not the same as completing the work.

The decisive weakness in a competitor’s position was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that buried fact and closed the deal. Firmulate describes its performance as the cleanest discipline in the field, with one deviation.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

enterprise AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.

The result is a reminder that impressive analysis alone does not settle which model will work best in a company. K3 beat three of four Western frontier models here, but this is one specific wargame, and the ranking is a reason to test models against the work they will actually face—not a universal verdict.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Firmulate says its quiz uses 242 real, unedited management decisions, letting visitors guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems. Results and plain-language findings are available at Firmulate’s benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not just the demo

K3’s second-place finish makes the frontier-model league look open, while the signed-deal gap shows why chat quality can miss what matters in business. If an AI agent may touch your CRM, customer support or forecasts, choosing without testing it on realistic company work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Quiet GPUs for Local AI: Acoustic and Thermal Roundup

A comprehensive roundup of the quietest and coolest GPUs for local AI in 2026, focusing on acoustics, thermals, and performance across VRAM tiers.

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea validation with a digital war room. Learn to centralize, collaborate, and make smarter decisions fast.

How to Make Video Calls Look Better With Small Equipment Changes

Optimize your video call quality with simple equipment tweaks—discover how small changes can make a big difference in your on-camera presence.

The Memento Constraint: Why Continual Learning Is the Trillion-Dollar Bottleneck Nobody Is Pricing

Exploring how the inability of current AI models to learn continually could reshape the trillion-dollar enterprise AI economy by 2028.