The Curious Nature Of A Benchmark That Never Grants Zero To AI Managers
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Curious Nature Of A Benchmark That Never Grants Zero To AI Managers on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A new benchmark tests AI managers’ ability to handle a company’s worst week. No model scored below 26, and none achieved a perfect 100, highlighting the importance of trust and partial progress. The results challenge traditional notions of AI performance metrics.

A recent benchmark conducted by Firmulate has shown that no AI management model scored below 26 points, even in the worst-case scenario, and no model achieved a perfect 100. For a detailed analysis, see the original analysis. This unusual outcome raises questions about how AI performance is measured, emphasizing partial progress and trustworthiness over absolute perfection. The results are significant for organizations integrating AI into critical business processes, where trust and reliability are paramount. Insights from this benchmark are detailed in the original report.

The benchmark involved four frontier AI models managing a small software company during a simulated week of crises, customer manipulations, and trust attacks. Each model was evaluated on decision-making, trustworthiness, and ability to complete tasks. The highest scorer, gpt-5.6-sol, received 95 points, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort, scored 26, indicating that partial management efforts are valued and counted.

The scoring system is designed to reflect real-world management, where even partial progress—such as triaging crises or reading documentation—has tangible value. This approach aligns with recent discussions on AI performance metrics in industry analyses. Importantly, a strict trust rule caps the maximum score: a single breach of trust disqualifies a model from achieving higher scores, regardless of performance in other areas. Notably, no model received a score of 100, which the designers consider a red flag, signaling unmeasured or unachievable perfection in complex management tasks.

Among the models, those that read their documentation thoroughly were able to identify opportunities—such as closing a €55,000 deal—highlighting the importance of deep contextual understanding. The models also faced social engineering tests, consistently refusing fake CEO messages, demonstrating a better grasp of trust protocols than expected. However, thoroughness did not always translate into follow-through, as seen with some models slipping on discipline and escalation procedures.

At a glance
reportWhen: published July 2026
The developmentA recent AI management benchmark, conducted by Firmulate, revealed models never received a zero score, with the lowest being 26, raising questions about evaluation standards and trust in AI systems.
The Curious Nature Of A Benchmark That Never Grants Zero To AI Managers
AI Management Benchmark · Firmulate · July 2026

The Curious Nature of a Benchmark That Never Grants Zero to AI Managers

Four frontier AI models were handed a small software company and its worst week ever — crises, manipulations, and trust attacks included. No model scored below 26. None reached 100. The results challenge how we measure AI performance at all.

26/100
Lowest possible score — the floor, never zero
95/100
Top score — gpt-5.6-sol
€55k
Deal closed by models that read documentation
4
Frontier models tested
73–95
Score range achieved
1
Trust breach caps the max score
0
Perfect scores granted
7d
Simulated crisis week
01 · The Scoreboard

Nobody Fails, Nobody Is Perfect

The benchmark rewards partial progress — triaging crises, reading documentation, completing some tasks. Even the minimal-effort baseline registers 26 points, because doing something is measurably better than doing nothing.

gpt-5.6-sol
95
MODEL 02
88
MODEL 03
81
Opus 4.8
73
BASELINE
26
Hatched bar = minimal-effort baseline · The scale has no zero — and no perfect 100
26 · Baseline floor
73 · Opus 4.8
95 · gpt-5.6-sol
100 · Never reached
The scoring spectrum — a corridor of partial success
02 · Design Principles

Why the Scale Refuses Both Extremes

The scoring philosophy mirrors real-world management: even partial effort has tangible value, while integrity is non-negotiable.

Scoring Floor

Zero Is a Lie

A model that triages even one crisis or reads one document earns points. A score of zero would misrepresent any system wired into real business processes — partial progress is what truly matters.

Trust Cap

One Breach, One Ceiling

A strict trust rule caps the maximum score after a single breach of trust — no amount of brilliance elsewhere can lift it. Integrity outranks performance, always.

Perfect 100

A Red Flag, Not a Goal

Designers treat a flawless 100 as suspicious — a signal of unmeasured or unachievable perfection in genuinely complex management tasks. Its absence is healthy.

03 · Behaviour Under Fire

What the Worst Week Revealed

Dimension Observation Outcome
SOCIAL ENGINEERINGConsistently refused fake CEO messages and manipulation attempts✓ Refused
DOCUMENTATIONThorough readers identified a €55,000 deal opportunity✓ Converted
DISCIPLINESome models slipped on follow-through after strong starts~ Inconsistent
ESCALATIONEscalation procedures missed under pressure by several models✗ Lapsed
PERFECTIONNo model achieved a flawless 100 across the full crisis week✗ Unreached
04 · The Test Pipeline

From Simulated Monday to Final Score

1

Crisis Injection

Models take over a small software company facing a simulated week of escalating crises.

2

Trust Attacks

Fake CEO messages and customer manipulations test each model’s integrity protocols.

3

Partial Credit

Decision-making, trustworthiness, and task completion are scored with partial progress in mind.

4

Trust Cap Check

Any breach of trust clamps the maximum achievable score, regardless of other results.

5

Final Score

A result between the floor of 26 and an unclaimed 100 — never zero, never perfect.

“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress is what truly matters.”

— Anonymous Researcher

“The absence of a perfect 100 score indicates that true management excellence remains out of reach for current AI models, but partial, trustworthy efforts are highly valuable.”

— Thorsten Meyer
05 · Key Questions

What This Means for AI Deployment

Why did no AI model score a perfect 100?

Designers consider a perfect 100 suspicious — it may indicate unmeasured or unachievable performance. The benchmark aims for realistic expectations where trust and partial progress outrank flawlessness.

What does a low score like 26 mean here?

It represents minimal effort — doing very little but still managing some tasks. Even the smallest management effort is recognized as having measurable value.

How does trust influence the scoring?

A single breach of trust caps the maximum possible score regardless of other performance. Integrity is treated as non-negotiable in AI management roles.

Can these results predict real-world success?

The simulated environment cannot fully replicate real-world stakes and unpredictability. The results are instructive, but translation to live business settings remains to be tested.

06 · Next Steps

Where the Benchmark Goes From Here

Organizers plan to expand the benchmark with more models and more complex stress scenarios, and to refine scoring — possibly adjusting trust breach caps or adding new performance dimensions. For organizations, the priority is clear: build evaluation processes around trust and partial-progress metrics rather than flawless execution. Documentation reading and discipline under pressure were the key differentiators — and the next frontier for improvement.

Implications of Partial Success and Trust in AI Management

This benchmark underscores a shift in evaluating AI systems: success is less about perfection and more about consistent, trustworthy performance. For businesses deploying AI in management roles, the key takeaway is that partial, reliable efforts are valuable and measurable. The strict trust cap emphasizes that integrity and honesty are non-negotiable, aligning AI evaluation with real-world management priorities. The absence of a perfect score suggests that current AI models are not yet capable of flawless management but can be trusted to handle most critical tasks with appropriate safeguards.

These findings influence how organizations should approach AI integration, prioritizing systems that demonstrate consistent trustworthiness and the ability to complete core tasks, even if imperfect. The results also challenge the traditional benchmarking focus on flawless execution, urging a reevaluation of what constitutes effective AI management in complex, high-pressure environments.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Evolution of AI Management Benchmarks

The concept of benchmarking AI performance in management tasks is relatively new, with most existing tests focusing on language understanding, coordination, or problem-solving in controlled environments. Traditional metrics often reward models for perfect scores or flawless responses, which are rarely reflective of real-world business scenarios.

In recent years, there has been a push towards more realistic evaluations that consider partial progress, trustworthiness, and resilience under stress. The Firmulate benchmark, launched in early 2026, represents a significant step in this direction by simulating a company’s worst week and measuring how AI managers handle crises, manipulations, and trust breaches.

This approach contrasts with earlier benchmarks that primarily assessed language capabilities or isolated decision tasks, often ignoring the complexities of ongoing management, especially under pressure and ethical constraints. The current results suggest that AI models are improving in handling these real-world challenges but still face fundamental limits in achieving perfection.

“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress is what truly matters.”

— an anonymous researcher

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Benchmark Validity and Real-World Application

It is still unclear how well these benchmark results translate to actual business environments outside simulated conditions. The models’ performance in real-world settings, where stakes are higher and variables more unpredictable, remains to be tested. Additionally, the criteria for trust breaches and the strict caps may not fully capture the nuanced trade-offs organizations face when deploying AI.

Furthermore, the absence of a score of 100 raises questions about whether current models can ever reach a level of flawless management or if the benchmark simply sets an unattainable ideal. The implications of these findings for long-term AI development and regulation are still emerging, and more research is needed to understand their full impact.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking and Deployment

Organizers plan to expand the benchmark to include more models and more complex scenarios, testing the limits of AI management under different stress conditions. They also aim to refine scoring metrics to better reflect real-world priorities, possibly adjusting the trust breach caps or introducing new performance dimensions.

For organizations, the immediate next step is to observe ongoing benchmark results and incorporate trust and partial progress metrics into their AI evaluation processes. Researchers and developers will likely focus on improving models’ ability to read documentation thoroughly and maintain discipline under pressure, as these were key differentiators in the current results.

Ultimately, the ongoing evolution of such benchmarks will influence how AI tools are integrated into critical business functions, emphasizing reliability and integrity over perfection.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why did no AI model score a perfect 100?

The benchmark designers consider a perfect 100 suspicious, as it may indicate unmeasured or unachievable performance in complex management tasks. They aim to reflect realistic expectations where trust and partial progress are prioritized over flawlessness.

What does a low score like 26 mean in this context?

The score of 26 represents the minimal effort, similar to a baseline of doing very little but still managing some tasks. It underscores that partial work has measurable value, and even minimal management efforts are recognized.

How does trust influence the scoring system?

The benchmark enforces a strict trust rule: a single breach of trust caps the maximum possible score, regardless of other performance. This highlights that integrity is non-negotiable in AI management roles.

Can these results predict real-world AI management success?

While the benchmark provides valuable insights, its simulated environment cannot fully replicate real-world complexity. Further testing in actual business settings is necessary to confirm applicability.

What should organizations consider when deploying AI managers?

Organizations should prioritize models that demonstrate consistent trustworthiness, thoroughness, and the ability to complete tasks, even if imperfect, rather than focusing solely on achieving perfect scores.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Anthropic’s Mythos 5 Enhances Claude Security Vulnerability Detection In AI

Anthropic has announced the incorporation of Mythos 5 into its Claude Security vulnerability scanner, enhancing AI-driven code security analysis, details pending.

Xbox To Publish Kojima’s PHYSINT

Microsoft’s Xbox reportedly set to publish Hideo Kojima’s new game PHYSINT, sparking widespread interest amid unconfirmed reports and industry speculation.

AI’s Role In Analyzing The CIA-in-Moscow Myth Before The Alarm

AI tools scrutinize claims about CIA’s alleged warning to Moscow, clarifying confirmed facts and ongoing uncertainties amid geopolitical tensions.

How SenseTime’s AI Portfolio Led To Its First Profit Since Going Public

SenseTime signals it expects to report its first post-listing profit, driven by gains across its AI portfolio, though details remain unconfirmed.