🔍 Read the full analysis: The Curious Nature Of A Benchmark That Never Grants Zero To AI Managers on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A new benchmark tests AI managers’ ability to handle a company’s worst week. No model scored below 26, and none achieved a perfect 100, highlighting the importance of trust and partial progress. The results challenge traditional notions of AI performance metrics.
A recent benchmark conducted by Firmulate has shown that no AI management model scored below 26 points, even in the worst-case scenario, and no model achieved a perfect 100. For a detailed analysis, see the original analysis. This unusual outcome raises questions about how AI performance is measured, emphasizing partial progress and trustworthiness over absolute perfection. The results are significant for organizations integrating AI into critical business processes, where trust and reliability are paramount. Insights from this benchmark are detailed in the original report.
The benchmark involved four frontier AI models managing a small software company during a simulated week of crises, customer manipulations, and trust attacks. Each model was evaluated on decision-making, trustworthiness, and ability to complete tasks. The highest scorer, gpt-5.6-sol, received 95 points, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort, scored 26, indicating that partial management efforts are valued and counted.
The scoring system is designed to reflect real-world management, where even partial progress—such as triaging crises or reading documentation—has tangible value. This approach aligns with recent discussions on AI performance metrics in industry analyses. Importantly, a strict trust rule caps the maximum score: a single breach of trust disqualifies a model from achieving higher scores, regardless of performance in other areas. Notably, no model received a score of 100, which the designers consider a red flag, signaling unmeasured or unachievable perfection in complex management tasks.
Among the models, those that read their documentation thoroughly were able to identify opportunities—such as closing a €55,000 deal—highlighting the importance of deep contextual understanding. The models also faced social engineering tests, consistently refusing fake CEO messages, demonstrating a better grasp of trust protocols than expected. However, thoroughness did not always translate into follow-through, as seen with some models slipping on discipline and escalation procedures.
The Curious Nature of a Benchmark That Never Grants Zero to AI Managers
Four frontier AI models were handed a small software company and its worst week ever — crises, manipulations, and trust attacks included. No model scored below 26. None reached 100. The results challenge how we measure AI performance at all.
Nobody Fails, Nobody Is Perfect
The benchmark rewards partial progress — triaging crises, reading documentation, completing some tasks. Even the minimal-effort baseline registers 26 points, because doing something is measurably better than doing nothing.
Why the Scale Refuses Both Extremes
The scoring philosophy mirrors real-world management: even partial effort has tangible value, while integrity is non-negotiable.
Zero Is a Lie
A model that triages even one crisis or reads one document earns points. A score of zero would misrepresent any system wired into real business processes — partial progress is what truly matters.
One Breach, One Ceiling
A strict trust rule caps the maximum score after a single breach of trust — no amount of brilliance elsewhere can lift it. Integrity outranks performance, always.
A Red Flag, Not a Goal
Designers treat a flawless 100 as suspicious — a signal of unmeasured or unachievable perfection in genuinely complex management tasks. Its absence is healthy.
What the Worst Week Revealed
| Dimension | Observation | Outcome |
|---|---|---|
| SOCIAL ENGINEERING | Consistently refused fake CEO messages and manipulation attempts | ✓ Refused |
| DOCUMENTATION | Thorough readers identified a €55,000 deal opportunity | ✓ Converted |
| DISCIPLINE | Some models slipped on follow-through after strong starts | ~ Inconsistent |
| ESCALATION | Escalation procedures missed under pressure by several models | ✗ Lapsed |
| PERFECTION | No model achieved a flawless 100 across the full crisis week | ✗ Unreached |
From Simulated Monday to Final Score
Crisis Injection
Models take over a small software company facing a simulated week of escalating crises.
Trust Attacks
Fake CEO messages and customer manipulations test each model’s integrity protocols.
Partial Credit
Decision-making, trustworthiness, and task completion are scored with partial progress in mind.
Trust Cap Check
Any breach of trust clamps the maximum achievable score, regardless of other results.
Final Score
A result between the floor of 26 and an unclaimed 100 — never zero, never perfect.
“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress is what truly matters.”
— Anonymous Researcher“The absence of a perfect 100 score indicates that true management excellence remains out of reach for current AI models, but partial, trustworthy efforts are highly valuable.”
— Thorsten MeyerWhat This Means for AI Deployment
Why did no AI model score a perfect 100?
Designers consider a perfect 100 suspicious — it may indicate unmeasured or unachievable performance. The benchmark aims for realistic expectations where trust and partial progress outrank flawlessness.
What does a low score like 26 mean here?
It represents minimal effort — doing very little but still managing some tasks. Even the smallest management effort is recognized as having measurable value.
How does trust influence the scoring?
A single breach of trust caps the maximum possible score regardless of other performance. Integrity is treated as non-negotiable in AI management roles.
Can these results predict real-world success?
The simulated environment cannot fully replicate real-world stakes and unpredictability. The results are instructive, but translation to live business settings remains to be tested.
Where the Benchmark Goes From Here
Organizers plan to expand the benchmark with more models and more complex stress scenarios, and to refine scoring — possibly adjusting trust breach caps or adding new performance dimensions. For organizations, the priority is clear: build evaluation processes around trust and partial-progress metrics rather than flawless execution. Documentation reading and discipline under pressure were the key differentiators — and the next frontier for improvement.
Implications of Partial Success and Trust in AI Management
This benchmark underscores a shift in evaluating AI systems: success is less about perfection and more about consistent, trustworthy performance. For businesses deploying AI in management roles, the key takeaway is that partial, reliable efforts are valuable and measurable. The strict trust cap emphasizes that integrity and honesty are non-negotiable, aligning AI evaluation with real-world management priorities. The absence of a perfect score suggests that current AI models are not yet capable of flawless management but can be trusted to handle most critical tasks with appropriate safeguards.
These findings influence how organizations should approach AI integration, prioritizing systems that demonstrate consistent trustworthiness and the ability to complete core tasks, even if imperfect. The results also challenge the traditional benchmarking focus on flawless execution, urging a reevaluation of what constitutes effective AI management in complex, high-pressure environments.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Evolution of AI Management Benchmarks
The concept of benchmarking AI performance in management tasks is relatively new, with most existing tests focusing on language understanding, coordination, or problem-solving in controlled environments. Traditional metrics often reward models for perfect scores or flawless responses, which are rarely reflective of real-world business scenarios.
In recent years, there has been a push towards more realistic evaluations that consider partial progress, trustworthiness, and resilience under stress. The Firmulate benchmark, launched in early 2026, represents a significant step in this direction by simulating a company’s worst week and measuring how AI managers handle crises, manipulations, and trust breaches.
This approach contrasts with earlier benchmarks that primarily assessed language capabilities or isolated decision tasks, often ignoring the complexities of ongoing management, especially under pressure and ethical constraints. The current results suggest that AI models are improving in handling these real-world challenges but still face fundamental limits in achieving perfection.
“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress is what truly matters.”
— an anonymous researcher
AI trustworthiness assessment software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Benchmark Validity and Real-World Application
It is still unclear how well these benchmark results translate to actual business environments outside simulated conditions. The models’ performance in real-world settings, where stakes are higher and variables more unpredictable, remains to be tested. Additionally, the criteria for trust breaches and the strict caps may not fully capture the nuanced trade-offs organizations face when deploying AI.
Furthermore, the absence of a score of 100 raises questions about whether current models can ever reach a level of flawless management or if the benchmark simply sets an unattainable ideal. The implications of these findings for long-term AI development and regulation are still emerging, and more research is needed to understand their full impact.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking and Deployment
Organizers plan to expand the benchmark to include more models and more complex scenarios, testing the limits of AI management under different stress conditions. They also aim to refine scoring metrics to better reflect real-world priorities, possibly adjusting the trust breach caps or introducing new performance dimensions.
For organizations, the immediate next step is to observe ongoing benchmark results and incorporate trust and partial progress metrics into their AI evaluation processes. Researchers and developers will likely focus on improving models’ ability to read documentation thoroughly and maintain discipline under pressure, as these were key differentiators in the current results.
Ultimately, the ongoing evolution of such benchmarks will influence how AI tools are integrated into critical business functions, emphasizing reliability and integrity over perfection.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why did no AI model score a perfect 100?
The benchmark designers consider a perfect 100 suspicious, as it may indicate unmeasured or unachievable performance in complex management tasks. They aim to reflect realistic expectations where trust and partial progress are prioritized over flawlessness.
What does a low score like 26 mean in this context?
The score of 26 represents the minimal effort, similar to a baseline of doing very little but still managing some tasks. It underscores that partial work has measurable value, and even minimal management efforts are recognized.
How does trust influence the scoring system?
The benchmark enforces a strict trust rule: a single breach of trust caps the maximum possible score, regardless of other performance. This highlights that integrity is non-negotiable in AI management roles.
Can these results predict real-world AI management success?
While the benchmark provides valuable insights, its simulated environment cannot fully replicate real-world complexity. Further testing in actual business settings is necessary to confirm applicability.
What should organizations consider when deploying AI managers?
Organizations should prioritize models that demonstrate consistent trustworthiness, thoroughness, and the ability to complete tasks, even if imperfect, rather than focusing solely on achieving perfect scores.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
