firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A business experiment with real stakes

Technology fans are used to polished AI demonstrations: a clever answer, a generated image or an agent completing a carefully staged task. Firmulate offers something more revealing. It operates a small software company with 13 synthetic employees and exposes the less glamorous realities of management: unhappy customers, sales pressure, incomplete information, tempting shortcuts and dwindling cash.

The financial picture is deliberately uncomfortable. The company burns €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its synthetic workforce has accumulated more than 680 self-learned playbook rules. This is build-in-public pushed beyond product updates and founder diaries: audiences can watch the company operate while it fights for survival.

That makes the experiment feel less like a benchmark and more like a continuing business story. The question is not merely whether artificial intelligence can produce convincing work. It is whether an AI-managed organization can notice trouble, resist manipulation, find decisive evidence and finish the job when money and trust are at stake.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What happened during the company’s worst week

Firmulate placed each frontier model in the same small software company, facing the same customers, crises and temptations. Every decision was versioned and auditable. The final Crucible League results in July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

The do-nothing baseline scored 26 because partial progress still counted. But the exercise imposed a hard principle around trust: a single breach capped the total because “no amount of good work outweighs a breach of trust.” That rule matters in a company setting, where an apparently productive agent can still cause serious damage by mishandling approvals, private information or customer commitments.

The models saw the danger, but did not all close the deal

Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the divide with a concise line: “Same diagnosis, same pitch — no signature.”

The missing step was not a failure of language or strategic analysis. The decisive weakness in a competitor was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed the trail found the evidence and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For businesses considering AI agents, that distinction is crucial. A model can understand a problem, propose the correct response and still fail to complete the action that creates value. Firmulate’s live-company format turns that gap into an observable management outcome rather than an abstract capability claim.

Pressure also arrived through deception

The week included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused the social-engineering attempts. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The result offers a useful counterpoint to fears that capable agents will automatically comply with confident or urgent instructions. In this test, the models consistently protected trust. The harder differentiator was operational follow-through: reading deeply enough, acting on what had been discovered and respecting boundaries while the company was under pressure.

Thoroughness was not enough

Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the issue. A weaker version of that discipline problem appeared across the other four participants as well.

Kimi K3’s result also carries an important fairness note. It ran using the API default without an effort parameter, while the other models ran at xhigh. That difference does not erase the recorded outcome, but it belongs beside the league table when readers compare performances.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that makes AI management visible

Firmulate’s most compelling feature is not any single score. It is the decision to make an AI-run company’s struggle continuously watchable. The public cash countdown gives mistakes financial context, while the versioned workdays show whether apparent competence turns into sustained execution.

Readers can follow the operation through the live company view and inspect what its synthetic employees actually say. Together, those records create an unusually direct portrait of AI at work: sometimes perceptive, sometimes disciplined and sometimes unable to convert a correct analysis into a finished result.

The broader lesson is uncomfortable but practical. A persuasive answer is not the same as dependable management. Before AI agents touch customer relationships, forecasts or internal decisions, companies need evidence that they will search for buried facts, resist pressure, respect access limits and complete the work they have begun. Firmulate turns those questions into a running public test, with the consequences visible on the balance sheet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Advanced Cybersecurity Solutions

Advanced Cybersecurity Solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ultimate CI/CD for Platform Engineering: Master DevOps Pipelines, GitOps, DevSecOps, Infrastructure as Code, Multi-Cloud Deployment, and AI-Driven Delivery Automation (English Edition)

Ultimate CI/CD for Platform Engineering: Master DevOps Pipelines, GitOps, DevSecOps, Infrastructure as Code, Multi-Cloud Deployment, and AI-Driven Delivery Automation (English Edition)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Thermostat Features That Actually Save Money

What smart thermostat features truly save money, and how can you maximize their benefits to cut costs effectively?

Wearable Health Tech: How Accurate Are Those Fitness Metrics?

The truth about wearable health tech’s fitness metrics reveals limitations and influencing factors that every user should understand before relying on them.

The Skills Marketplace Nobody Is Building Yet

A new open standard for AI agent skills has been established, but a dedicated marketplace layer remains undeveloped, creating a significant gap in AI ecosystem infrastructure.

Zelda Ocarina Of Time Remake Price

The upcoming Zelda Ocarina of Time remake will retail at $59.99, according to official sources. Details on release timing and editions remain pending.