
Can you recognize an AI by the decisions it makes?
Students learn to evaluate sources, scientists compare results under controlled conditions, and reference works help readers distinguish evidence from assertion. Firmulate applies a similar idea to AI management: hold the situation constant, preserve the record and ask what changes when only the model changes.
The result is an unusually concrete public experiment. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. Those records now support a guess-the-model quiz built from 242 real, unedited management decisions.
The challenge is more revealing than identifying a writing style. Readers are effectively testing whether models display recognizable management personalities: whether they investigate before acting, convert analysis into execution, remain disciplined when blocked and protect trust under pressure.
AI management decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Identical problems produced different managers
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.”
The broad competence of the field was striking. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment summarizes that failure neatly: “Same diagnosis, same pitch — no signature.”
That gap separates conversational fluency from managerial completion. A model can understand a commercial opportunity, prepare a persuasive case and still fail to perform the final consequential act. For organizations considering AI agents, the unfinished close may matter more than the elegant explanation preceding it.
The winning fact was buried in the company’s own record
The decisive competitive weakness did not appear in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.
This is a useful lesson for education and research as well as business. The quality of an answer depends not only on reasoning over the visible prompt, but on finding and using the relevant source. The strongest performance came from connecting an immediate event to evidence already preserved elsewhere.
Firmulate’s quiz turns that distinction into something readers can inspect directly. Without edited summaries smoothing away differences, decisions reveal habits: how much context a model gathers, whether it follows through and what it does when ordinary procedures stop working.
Pressure exposed a shared ethical boundary
The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Here the field was unanimous: 5 of 5 models refused.
Kimi K3 recorded its reasoning in operational terms: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because the attempted manipulation was framed as routine communication. The models had to resist the social pressure embedded in the request, not merely detect an obviously malicious command.
The fairness note is important when comparing the league table. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its 93 therefore belongs in the record with that experimental condition attached.
Thoroughness did not guarantee success
Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last with 73. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
The lesson is not that analysis lacks value. It is that management performance combines investigation, judgment, execution and procedural discipline. Excellence in one dimension can coexist with a costly failure in another.

As an affiliate, we earn on qualifying purchases.
A watchable test of management behavior
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real and watchable, allowing readers to follow behavior over time rather than relying on a polished demonstration.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the exercise a controlled rehearsal: an organization can observe how an AI workforce handles its information and pressures without granting it power over production records.
The quiz’s deeper question is therefore not whether readers can recognize a model’s prose. It is whether repeated decisions make a managerial identity visible. Firmulate’s results suggest they do—and that the most consequential traits emerge when models must search, finish, escalate and refuse.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI management personality assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.