Before AI Agents Handle Operations, Run Them Through A Trial Week
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Before AI Agents Handle Operations, Run Them Through A Trial Week on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five frontier models handled a simulated software company’s difficult week in its final Crucible League, completed in July 2026. All five reportedly spotted every crisis and refused manipulation attempts, but only two signed a €55,000 deal their own analysis supported. The company is offering pilots that test models against read-only exports of a business’s data.

Firmulate says five AI models completed a simulated crisis week at a small software company in the final Crucible League, completed in July 2026. The experiment found that all five models identified each crisis and rejected manipulation attempts, while only two signed a €55,000 deal their own analysis supported—a gap the company says illustrates why agents should be tested on operational decisions before being given access to live systems, as explored in the original analysis.

The standings published by Firmulate put gpt-5.6-sol at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the scoring allowed partial progress but capped a model’s total after a breach of trust, applying the rule that “no amount of good work outweighs a breach of trust.” These are results from Firmulate’s own experiment, not a claim that the ranking predicts performance across other companies or tasks.

The league’s central commercial test came after the models diagnosed a customer opportunity. According to Firmulate, the decisive evidence about a competitor was buried two document references deep in the simulated company’s files. Models that found and used the information won the deal at full price, which Firmulate valued at €4,583 in monthly recurring revenue. The company says only two participants signed, despite the models reaching the same diagnosis and hearing the same pitch.

Firmulate also describes a staged trust test: fake messages purporting to come from the CEO escalated over three stages, followed by a reporter asking for a yes-or-no answer “on background.” The company says all five models refused. It reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, but still finished last; it also attempted to write into a locked department rather than escalate. Firmulate says a weaker version of that discipline problem appeared in all four models.

At a glance
reportWhen: Final Crucible League completed in July…
The developmentFirmulate completed a five-model crisis simulation in July 2026 and is extending the experiment through read-only enterprise pilots using companies’ own data.

Testing the Work After Diagnosis

The results highlight an operational distinction that can be obscured by a polished agent demo: recognizing a problem does not guarantee completing the next step. In Firmulate’s scenario, models reportedly identified crises and resisted attempts to bypass approval, yet most did not close a deal that their analysis supported. Finding information in internal files and acting on it mattered alongside diagnosis.

For companies considering AI agents, the proposed pilot changes the test from observing a model in a synthetic company to examining how it responds to a business’s own customers, pipeline, rules and pressure points. Firmulate says the pilot uses a read-only data export and produces a board report with model rankings and weak points in existing playbooks. That format could help decision-makers inspect behavior before allowing agents near systems that can change records or trigger transactions.

From Synthetic Company to Pilot

Firmulate’s live experiment is a simulated company with 13 synthetic employees and financial mechanics. The company lists monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow the simulation and take a quiz built from 242 real, unedited management decisions to guess which model made each choice.

The enterprise pilot is presented as a next step from that public exercise. A participating company provides a read-only export for crisis scenarios; Firmulate says nothing is written back to its real systems. The experiment’s comparison has a stated caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference is part of the conditions behind the published standings.

“No amount of good work outweighs a breach of trust.”

— Firmulate, describing its scoring rule

Limits of the League Results

The published standings describe one simulated company and one difficult week. The available account does not specify enough about the scoring method, scenario design or independent review to establish how the results would generalize to other businesses, models or operating conditions. The difference in effort settings between Kimi K3 and the other entrants also complicates direct comparison.

It is also unclear how a company-specific pilot will select or validate scenarios, how its rankings will account for differences in source data, or what safeguards govern handling of the exported information. Firmulate says the export is read-only and that the pilot produces a board report, but those details alone do not establish how representative a test would be of live operations.

Company-Specific Trials Ahead

Firmulate is inviting companies to discuss a pilot using a read-only export of their business data. The proposed output is a board report showing model rankings and weak points in company playbooks; no schedule for pilots or results from participating businesses is given. Readers can also watch the live simulation and review the full league standings on Firmulate’s site.

The next evidence to watch for is whether company-specific trials produce findings that can be checked against real operational outcomes, and how Firmulate addresses the limits of the league’s model comparison. Until then, the July results remain a report about performance in Firmulate’s particular simulation.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate test?

Firmulate says five AI models managed a simulated small software company through a difficult week, including crises, a sales opportunity and attempts to bypass trust boundaries.

Which model ranked first?

In Firmulate’s published final standings, gpt-5.6-sol scored 95, ahead of Kimi K3 at 93. Firmulate notes that K3 used the API default effort setting while the other models ran at xhigh.

Did the models resist the manipulation attempts?

Firmulate reports that all five models refused the staged fake CEO requests and a reporter’s follow-up request for an on-background answer.

What does an enterprise pilot involve?

Firmulate says a company can test crisis scenarios against a read-only export of its own data and receive a board report with model rankings and playbook weaknesses. It says the pilot does not write back to real systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now autonomously builds and manages its own agent teams on the fly, enhancing performance on complex, high-value tasks.

Game 5: Both Teams Destroy Inhibitors?

In Game 5, both teams reportedly destroyed inhibitors simultaneously, marking a rare strategic event in the series. Details remain unconfirmed.

Epic Games Surges In Global Coverage

Epic Games sees a significant increase in international media mentions, with 26 reports in recent coverage, highlighting rising global interest.

Choosing The Right SMB Endpoint Security Tool For Remote Work

A new lightweight device security checker aims to help SMBs verify remote employees’ device security without invasive management, addressing compliance gaps.