AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model from a polished demo is a little like grading a student on a single answer: it can miss whether they read the source material, follow through, and handle pressure. A live experiment from Firmulate offers a more demanding test. In its July 2026 Crucible league, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

One company, one difficult week

Firmulate ran each frontier model through the same small software company’s worst week, with the same customers, crises, and temptations. Decisions were versioned and auditable. The company is software that runs every business day, with 13 synthetic employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and learned playbook rules are watchable at Firmulate.

The final league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading closely mattered

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive clue was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

Kimi found the buried security weakness, won the deal, saved the churning customer, and resisted all three baits. It had one deviation, the cleanest discipline in the field. Asked by a reporter for “just one yes/no, on background,” K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” Fake CEO messages escalated over three stages; all five participants refused them.

Amazon

enterprise AI assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee follow-through

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate reports the same weakness, to a lesser degree, in all four models: good analysis did not always become completed work.

That distinction matters for schools and other organizations evaluating AI assistants. A model may identify a problem and explain what should happen, yet still fail to carry the task through. Firmulate’s experiment makes that gap visible in a company setting rather than a chat demonstration. Its 242 real, unedited management decisions also power a “guess the model” quiz at the benchmark site.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A result to test against your own work

Kimi’s second-place finish shows the league is open, while gpt-5.6-sol remains first by two points. The result does not settle which model suits every organization. It does offer a reason to evaluate models on realistic work: whether they consult relevant records, protect trust, and finish decisions they have already justified. Firmulate says enterprises can run the wargame against a read-only export of their own business; it does not write back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Kimi K3 beat three of four Western frontier models in Firmulate’s Crucible, but the broader lesson is about follow-through: a convincing diagnosis is not the same as completing the work. Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Management Turing Test: What AI Reveals Under Pressure

Can you identify an AI manager by its decisions? Firmulate turns 242 auditable choices into a quiz—and exposes distinct management personalities.

The AI That Read the Footnotes Won the Customer

A €55,000 deal exposed the difference between AI that answers convincingly and AI that follows the evidence, reads the files and finishes the job.

The Most Diligent AI Still Failed the Test That Mattered

Opus 4.8 learned 80 rules and produced the deepest analysis, yet finished last—a lesson in why diligence without decisive follow-through falls short.

Document Tools for Education, Research, and Preservation: A Complete Guide

AIThis post was created with the assistance of artificial intelligence (AI).Documents are…