AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

When You Grade AI Like a Manager, Not a Chatbot

We are used to judging artificial intelligence by how eloquently it answers. But the interesting questions start when an AI has to act — when it inherits customers, invoices, a crisis calendar and the temptation to cut corners, and every decision lands somewhere real. That is precisely the setup of Firmulate, a public experiment that runs frontier AI models as complete companies: same firm, same worst week, same temptations — only the model changes. For anyone interested in how we might someday teach or assess AI judgment the way we assess human managers, it is a fascinating live classroom.

The experiment is watchable in real time: a synthetic software company staffed by 13 AI employees, with real money mechanics — a burn rate of €105,000 per month against €2,300 in monthly recurring revenue — and a public cash countdown. The company has been running for 1,947 days, has self-learned 680+ playbook rules, and every workday is versioned so any decision can be replayed like a lab notebook entry.

Same Company, Same Crises, Different Brains

In the final Crucible League standings (July 2026), five models each ran the identical small software company through its worst week. The results: gpt-5.6-sol took first place with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline — essentially letting the company drift — scored just 26, which gives the numbers meaning: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Everyone Diagnosed the Disease; Two Prescribed the Cure

The headline finding is deceptively simple. All models spotted every crisis and refused every manipulation attempt. But only two of them actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap between competence and completion is invisible in chat demos, and it is exactly the kind of thing you only discover by running an AI under load.

The buried fact is even more instructive: the decisive competitor weakness that would win the deal sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In other words, the difference between first and fourth place wasn’t intelligence. It was diligence.

Pressure Tests and Social Engineering

The week included staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning reads like a textbook security policy: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Paradox

Opus 4.8 is the study’s most poignant profile: the most thorough participant, with +80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped (it attempted writes into a locked department instead of escalating). Notably, the same weakness appeared, weaker, in all four other models. Effort is not the same as judgment.

One methodological footnote deserves mention, in the spirit of good science: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second. Transparency about conditions is part of what makes the benchmark credible.

A Quiz Built From Real Decisions

For readers who want to test their own eye, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. It is a surprisingly humbling exercise — and a novel form of public assessment material.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Acting: Wargaming Your Own Business

Here is the natural next step, and it is open to enterprises now. Firmulate’s pilot lets a company run the same kind of wargame against a read-only export of its own business — your customers, your pipeline, your rules. Crisis scenarios (churn waves, price increases, competitor attacks, PR crises, social-engineering pressure) are played out against your own company, and the output is a board report with a model ranking and a map of the weak points in your own playbooks.

The critical guarantee: nothing ever writes back to real systems. It is a flight simulator, not an autopilot — you learn where your organization breaks before reality tests it for you.

If the experiment proves one thing, it is this: AI models differ less in what they know than in what they finish. Finding out which models complete the job on your data, under your crises, is now something you can measure rather than guess. Ready to run the wargame against your own company? Start your pilot at firmulate.com/pilot.html, or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Live AI Company Test Puts the Model League Up for Grabs

Kimi K3 placed second in Firmulate’s live AI company trial, ahead of three frontier models. The test suggests model choice needs evidence from real work.

Document Tools for Education, Research, and Preservation: A Complete Guide

AIThis post was created with the assistance of artificial intelligence (AI).Documents are…

The AI Test That Matters Begins After the Right Answer

Coding tests show whether AI can answer. Firmulate asks the harder question: can an agent manage under pressure without betraying trust?

Best Educational Science Kits For Kids Compared

Compare Snap Circuits and National Geographic science kits by age, learning style, replay value, price, and which is the better fit for your child.