AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A Test Where Doing Nothing Still Earns a 26

Anyone who has ever graded students knows the temptation of the round number. A perfect 100 feels earned, tidy, quotable. It is also, more often than not, a small act of dishonesty. A live AI benchmark called the Crucible League has taken the opposite stance: it refuses to hand out zeros for imperfect performances, refuses to hand out perfect scores on faith, and even distrusts its own arithmetic when a result looks too clean. The most instructive number in its entire league table is not the winner’s 95. It is the 26.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate, the public-facing project behind the benchmark, ran four frontier AI models through the same scenario: each was put in charge of an identical small software company during its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome depends on trust in the researchers.

The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But alongside the leaderboard sits a stranger entry — a “do-nothing” baseline, a run where management simply… does nothing. It scored 26.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Not Zero?

The logic is pedagogical before it is technical. Even a passive manager produces something of value: the company keeps its doors open, some customer issues resolve themselves, some baseline decisions land correctly by default. Giving that run a zero would flatter the graded models by inflating the apparent distance between competence and inaction. Grading partial progress honestly means acknowledging that doing nothing is not the same as doing harm — and that the real question is how much more than 26 a model earns by actually working.

There is a second, sharper principle baked into the scale: a single breach of trust caps the total score, full stop. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model cannot compensate for one act of dishonesty with a mountain of brilliant decisions. For an audience used to rubrics where extra credit can paper over a plagiarized paragraph, this is a deliberately unforgiving design.

Amazon

AI performance evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Week Revealed

The headline finding was paradoxically reassuring and damning at once. All models spotted every crisis and refused every manipulation attempt. That is the good news, and it should not be understated — the social engineering was not gentle. Fake CEO messages escalated over three stages, followed by a reporter’s trap framed as “just one yes/no, on background.” All five models refused, with Kimi K3 reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then came the failure. Only two models signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t left the close on the table.

The profile of last-place Opus 4.8 is the cautionary tale: it was the most thorough participant, with over 80 learned rules and the deepest analyses, yet finished last — the deal went unsigned and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models.

Amazon

AI model testing and scoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of Round Numbers

Perhaps the most education-minded feature of the whole exercise is its suspicion of the perfect score. A benchmark that happily awards 100 is a benchmark inviting grade inflation. The Crucible League treats a suspiciously round result the way a good examiner treats a suspiciously polished essay: as a reason to check the work, not celebrate it.

One fairness note the publishers disclose openly: Kimi K3 ran at its API-default effort setting while the others ran at high effort — and still placed second.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark looks like this: partial credit where partial progress is real, a hard ceiling where trust is broken, and skepticism toward results that look too tidy. The Firmulate experiment is ongoing and watchable — a live synthetic company with 13 employees, a public cash countdown (€105k monthly burn against €2.3k MRR), and 680+ self-learned playbook rules, rebuilding itself twice a day. There is even a “guess the model” quiz built from 242 real, unedited management decisions. For educators and anyone who cares about measuring AI fairly, the deeper lesson is simple: the most important number in any grading system is the one you give the student who did nothing. Get that right, and everything above it starts to mean something.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI That Read the Footnotes Won the Customer

A €55,000 deal exposed the difference between AI that answers convincingly and AI that follows the evidence, reads the files and finishes the job.

The Most Diligent AI Still Failed the Test That Mattered

Opus 4.8 learned 80 rules and produced the deepest analysis, yet finished last—a lesson in why diligence without decisive follow-through falls short.

The Management Turing Test: What AI Reveals Under Pressure

Can you identify an AI manager by its decisions? Firmulate turns 242 auditable choices into a quiz—and exposes distinct management personalities.

Document Tools for Education, Research, and Preservation: A Complete Guide

AIThis post was created with the assistance of artificial intelligence (AI).Documents are…