
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
When You Grade AI Like a Manager, Not a Chatbot
We are used to judging artificial intelligence by how eloquently it answers. But the interesting questions start when an AI has to act — when it inherits customers, invoices, a crisis calendar and the temptation to cut corners, and every decision lands somewhere real. That is precisely the setup of Firmulate, a public experiment that runs frontier AI models as complete companies: same firm, same worst week, same temptations — only the model changes. For anyone interested in how we might someday teach or assess AI judgment the way we assess human managers, it is a fascinating live classroom.
The experiment is watchable in real time: a synthetic software company staffed by 13 AI employees, with real money mechanics — a burn rate of €105,000 per month against €2,300 in monthly recurring revenue — and a public cash countdown. The company has been running for 1,947 days, has self-learned 680+ playbook rules, and every workday is versioned so any decision can be replayed like a lab notebook entry.
Same Company, Same Crises, Different Brains
In the final Crucible League standings (July 2026), five models each ran the identical small software company through its worst week. The results: gpt-5.6-sol took first place with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline — essentially letting the company drift — scored just 26, which gives the numbers meaning: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
Everyone Diagnosed the Disease; Two Prescribed the Cure
The headline finding is deceptively simple. All models spotted every crisis and refused every manipulation attempt. But only two of them actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap between competence and completion is invisible in chat demos, and it is exactly the kind of thing you only discover by running an AI under load.
The buried fact is even more instructive: the decisive competitor weakness that would win the deal sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In other words, the difference between first and fourth place wasn’t intelligence. It was diligence.
Pressure Tests and Social Engineering
The week included staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning reads like a textbook security policy: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Thoroughness Paradox
Opus 4.8 is the study’s most poignant profile: the most thorough participant, with +80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped (it attempted writes into a locked department instead of escalating). Notably, the same weakness appeared, weaker, in all four other models. Effort is not the same as judgment.
One methodological footnote deserves mention, in the spirit of good science: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second. Transparency about conditions is part of what makes the benchmark credible.
A Quiz Built From Real Decisions
For readers who want to test their own eye, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. It is a surprisingly humbling exercise — and a novel form of public assessment material.

From Watching to Acting: Wargaming Your Own Business
Here is the natural next step, and it is open to enterprises now. Firmulate’s pilot lets a company run the same kind of wargame against a read-only export of its own business — your customers, your pipeline, your rules. Crisis scenarios (churn waves, price increases, competitor attacks, PR crises, social-engineering pressure) are played out against your own company, and the output is a board report with a model ranking and a map of the weak points in your own playbooks.
The critical guarantee: nothing ever writes back to real systems. It is a flight simulator, not an autopilot — you learn where your organization breaks before reality tests it for you.
If the experiment proves one thing, it is this: AI models differ less in what they know than in what they finish. Finding out which models complete the job on your data, under your crises, is now something you can measure rather than guess. Ready to run the wargame against your own company? Start your pilot at firmulate.com/pilot.html, or write to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
