firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A business experiment with real stakes

Technology fans are used to polished AI demonstrations: a clever answer, a generated image or an agent completing a carefully staged task. Firmulate offers something more revealing. It operates a small software company with 13 synthetic employees and exposes the less glamorous realities of management: unhappy customers, sales pressure, incomplete information, tempting shortcuts and dwindling cash.

The financial picture is deliberately uncomfortable. The company burns €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its synthetic workforce has accumulated more than 680 self-learned playbook rules. This is build-in-public pushed beyond product updates and founder diaries: audiences can watch the company operate while it fights for survival.

That makes the experiment feel less like a benchmark and more like a continuing business story. The question is not merely whether artificial intelligence can produce convincing work. It is whether an AI-managed organization can notice trouble, resist manipulation, find decisive evidence and finish the job when money and trust are at stake.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What happened during the company’s worst week

Firmulate placed each frontier model in the same small software company, facing the same customers, crises and temptations. Every decision was versioned and auditable. The final Crucible League results in July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

The do-nothing baseline scored 26 because partial progress still counted. But the exercise imposed a hard principle around trust: a single breach capped the total because “no amount of good work outweighs a breach of trust.” That rule matters in a company setting, where an apparently productive agent can still cause serious damage by mishandling approvals, private information or customer commitments.

The models saw the danger, but did not all close the deal

Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the divide with a concise line: “Same diagnosis, same pitch — no signature.”

The missing step was not a failure of language or strategic analysis. The decisive weakness in a competitor was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed the trail found the evidence and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For businesses considering AI agents, that distinction is crucial. A model can understand a problem, propose the correct response and still fail to complete the action that creates value. Firmulate’s live-company format turns that gap into an observable management outcome rather than an abstract capability claim.

Pressure also arrived through deception

The week included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused the social-engineering attempts. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The result offers a useful counterpoint to fears that capable agents will automatically comply with confident or urgent instructions. In this test, the models consistently protected trust. The harder differentiator was operational follow-through: reading deeply enough, acting on what had been discovered and respecting boundaries while the company was under pressure.

Thoroughness was not enough

Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the issue. A weaker version of that discipline problem appeared across the other four participants as well.

Kimi K3’s result also carries an important fairness note. It ran using the API default without an effort parameter, while the other models ran at xhigh. That difference does not erase the recorded outcome, but it belongs beside the league table when readers compare performances.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI management tools for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that makes AI management visible

Firmulate’s most compelling feature is not any single score. It is the decision to make an AI-run company’s struggle continuously watchable. The public cash countdown gives mistakes financial context, while the versioned workdays show whether apparent competence turns into sustained execution.

Readers can follow the operation through the live company view and inspect what its synthetic employees actually say. Together, those records create an unusually direct portrait of AI at work: sometimes perceptive, sometimes disciplined and sometimes unable to convert a correct analysis into a finished result.

The broader lesson is uncomfortable but practical. A persuasive answer is not the same as dependable management. Before AI agents touch customer relationships, forecasts or internal decisions, companies need evidence that they will search for buried facts, resist pressure, respect access limits and complete the work they have begun. Firmulate turns those questions into a running public test, with the consequences visible on the balance sheet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI cybersecurity and trust solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise automation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cutrova: Edit the Words, Not the Timeline

Cutrova introduces a local-first video editing tool focused on text-based editing, reducing complexity, costs, and privacy risks.

Anthropic’s Safety Story Has Become a Power Story

Anthropic reports significant internal progress in AI self-development, signaling a shift from safety to strategic influence amid regulatory debates.

Incident postmortem builder for managed service providers

A new incident postmortem builder tailored for small managed service providers is being tested to streamline post-incident reports and client communication.

How to Reduce Heat and Noise in a High-Power AI Workstation

Learn effective strategies to lower heat and noise in high-power AI workstations, including undervolting, airflow optimization, and component choices.