
Choosing an AI model from a polished demo is a little like grading a student on a single answer: it can miss whether they read the source material, follow through, and handle pressure. A live experiment from Firmulate offers a more demanding test. In its July 2026 Crucible league, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
One company, one difficult week
Firmulate ran each frontier model through the same small software company’s worst week, with the same customers, crises, and temptations. Decisions were versioned and auditable. The company is software that runs every business day, with 13 synthetic employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and learned playbook rules are watchable at Firmulate.
The final league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Reading closely mattered
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive clue was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
Kimi found the buried security weakness, won the deal, saved the churning customer, and resisted all three baits. It had one deviation, the cleanest discipline in the field. Asked by a reporter for “just one yes/no, on background,” K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” Fake CEO messages escalated over three stages; all five participants refused them.
As an affiliate, we earn on qualifying purchases.
Thoroughness did not guarantee follow-through
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate reports the same weakness, to a lesser degree, in all four models: good analysis did not always become completed work.
That distinction matters for schools and other organizations evaluating AI assistants. A model may identify a problem and explain what should happen, yet still fail to carry the task through. Firmulate’s experiment makes that gap visible in a company setting rather than a chat demonstration. Its 242 real, unedited management decisions also power a “guess the model” quiz at the benchmark site.
As an affiliate, we earn on qualifying purchases.
A result to test against your own work
Kimi’s second-place finish shows the league is open, while gpt-5.6-sol remains first by two points. The result does not settle which model suits every organization. It does offer a reason to evaluate models on realistic work: whether they consult relevant records, protect trust, and finish decisions they have already justified. Firmulate says enterprises can run the wargame against a read-only export of their own business; it does not write back to real systems.

Kimi K3 beat three of four Western frontier models in Firmulate’s Crucible, but the broader lesson is about follow-through: a convincing diagnosis is not the same as completing the work. Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
