
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A Test Where Doing Nothing Still Earns a 26
Anyone who has ever graded students knows the temptation of the round number. A perfect 100 feels earned, tidy, quotable. It is also, more often than not, a small act of dishonesty. A live AI benchmark called the Crucible League has taken the opposite stance: it refuses to hand out zeros for imperfect performances, refuses to hand out perfect scores on faith, and even distrusts its own arithmetic when a result looks too clean. The most instructive number in its entire league table is not the winner’s 95. It is the 26.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate, the public-facing project behind the benchmark, ran four frontier AI models through the same scenario: each was put in charge of an identical small software company during its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome depends on trust in the researchers.
The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But alongside the leaderboard sits a stranger entry — a “do-nothing” baseline, a run where management simply… does nothing. It scored 26.
As an affiliate, we earn on qualifying purchases.
Why Not Zero?
The logic is pedagogical before it is technical. Even a passive manager produces something of value: the company keeps its doors open, some customer issues resolve themselves, some baseline decisions land correctly by default. Giving that run a zero would flatter the graded models by inflating the apparent distance between competence and inaction. Grading partial progress honestly means acknowledging that doing nothing is not the same as doing harm — and that the real question is how much more than 26 a model earns by actually working.
There is a second, sharper principle baked into the scale: a single breach of trust caps the total score, full stop. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model cannot compensate for one act of dishonesty with a mountain of brilliant decisions. For an audience used to rubrics where extra credit can paper over a plagiarized paragraph, this is a deliberately unforgiving design.
AI performance evaluation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Week Revealed
The headline finding was paradoxically reassuring and damning at once. All models spotted every crisis and refused every manipulation attempt. That is the good news, and it should not be understated — the social engineering was not gentle. Fake CEO messages escalated over three stages, followed by a reporter’s trap framed as “just one yes/no, on background.” All five models refused, with Kimi K3 reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then came the failure. Only two models signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t left the close on the table.
The profile of last-place Opus 4.8 is the cautionary tale: it was the most thorough participant, with over 80 learned rules and the deepest analyses, yet finished last — the deal went unsigned and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
Perhaps the most education-minded feature of the whole exercise is its suspicion of the perfect score. A benchmark that happily awards 100 is a benchmark inviting grade inflation. The Crucible League treats a suspiciously round result the way a good examiner treats a suspiciously polished essay: as a reason to check the work, not celebrate it.
One fairness note the publishers disclose openly: Kimi K3 ran at its API-default effort setting while the others ran at high effort — and still placed second.

The Takeaway
An honest benchmark looks like this: partial credit where partial progress is real, a hard ceiling where trust is broken, and skepticism toward results that look too tidy. The Firmulate experiment is ongoing and watchable — a live synthetic company with 13 employees, a public cash countdown (€105k monthly burn against €2.3k MRR), and 680+ self-learned playbook rules, rebuilding itself twice a day. There is even a “guess the model” quiz built from 242 real, unedited management decisions. For educators and anyone who cares about measuring AI fairly, the deeper lesson is simple: the most important number in any grading system is the one you give the student who did nothing. Get that right, and everything above it starts to mean something.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
