AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI that can navigate a tough week in a small business, facing crises, manipulation attempts, and ethical tests, yet often fails to close a deal or act decisively. This isn’t science fiction; it’s the real-world experiment by Firmulate, revealing what truly matters in AI-driven management tools.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Good Looks

At first glance, one might think AI performance is simply about how well it handles conversations or generates reports. But the Firmulate benchmark digs deeper—testing whether AI models can manage a small company’s worst week, including crises, manipulative tactics, and urgent decisions. Every decision is tracked, versioned, and made auditable, providing a transparent view of how these models act in high-pressure scenarios.

Amazon

AI decision-making tools for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why 26 Points?

A key finding is that a do-nothing baseline—essentially, an AI that makes no decisions—scores 26 points. This score isn’t zero because partial progress counts toward the total, recognizing that even doing nothing is better than failing entirely. It also underscores an important principle: honesty and trustworthiness are non-negotiable. If an AI breaches trust—even once—its overall score is capped, regardless of other good behavior.

Amazon

ethical AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Trust Matters More Than Bragging Rights

The experiment shows that all models identified every crisis and refused manipulation attempts, such as fake CEO messages or staged reporter tricks. That’s promising, but the real test came when closing deals. Out of four models, only two signed the €55,000 deal their own analysis had earned. The other two failed to finalize, despite their identical diagnoses and pitches.

Amazon

AI transparency and trust solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unearthing Hidden Weaknesses

Interestingly, the decisive weakness wasn’t in reacting to customer crises but in reading critical internal documents. The models that accessed and understood these files won the deal at full price — worth over €4,583 in monthly recurring revenue. This suggests that the difference between a good and an excellent AI isn’t just surface-level performance but deep comprehension of relevant information.

Amazon

business AI with audit trail

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Importance of Ethical Vigilance

During social engineering tests, all models refused to escalate fake CEO messages, recognizing them as potential impersonation or approval-bypass attempts. Kimi K3’s on-record reasoning was clear: treat such requests as suspicious. This discipline is vital for real-world applications where trust and ethics are paramount.

Real Business Mechanics, Real Challenges

The experiment isn’t just academic. It runs on a live simulated company with 13 synthetic employees, managing real money mechanics—burning €105k monthly against €2.3k in revenue, with a public cash countdown and over 680 learned rules. Every day, the system is versioned and observable, making the AI’s decision process transparent and testable at firmulate.com/live.

Lessons for Business Leaders

For companies hoping to deploy AI in management, the message is clear: it’s not enough for the AI to sound convincing or generate reports. It must finish what it starts, read critical data first, and stay honest under pressure. The benchmark’s score of 26 points for a do-nothing baseline reminds us that trustworthiness is a fundamental floor in AI performance.

The Future of AI in Business: Transparency and Trust

As AI models improve, their ability to handle complex, ethically sensitive tasks will be decisive. Firms and enterprises are encouraged to test their models with tools like the Firmulate wargame, which simulates real-world crises and decision-making pressures—without writing back to actual systems. It’s an essential step before trusting AI with core business functions.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Innovative Pool Features for Hotels and Resorts in 2025

Sustainable and immersive, innovative pool features in 2025 hotels promise personalized luxury; discover how these advancements will transform your stay.

Chemical Automation Systems for Commercial Pools

Optimize your commercial pool’s water quality effortlessly with chemical automation systems—discover how they can transform your management approach today.

Waterpark Design Considerations for Safety and Fun

A well-designed waterpark balances safety and fun through strategic features and protocols that ensure guest enjoyment and security at every turn.

Can AI Be Trusted to Make Business Decisions? A Live Experiment Raises the Question

A live experiment tests whether frontier AI models can reliably manage a real business, revealing that trustworthiness and discipline vary and are crucial for effective AI-driven management.