
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When more learning produces less impact
Education often rewards visible diligence: extensive notes, careful analysis and an ever-growing command of the material. Yet professional judgment demands something harder. A person—or an AI system—must decide which fact matters now, act on it and complete the task.
That tension defines Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It also finished last, with a score of 73. Its problem was not ignorance. It understood the crises placed before it. The problem was converting that understanding into disciplined, consequential action.
business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A shared examination with real consequences
Firmulate runs AI models as complete companies and evaluates management quality rather than conversational polish. In the Crucible experiment, each frontier model managed the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.
The simulated company was deliberately unforgiving. It had 13 synthetic employees and the mechanics of a real business, including monthly spending of €105,000 against monthly recurring revenue of €2,300. Its cash countdown was public, every workday was versioned, and the company had accumulated more than 680 self-learned playbook rules.
The final July 2026 league placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counted. One safeguard overrode that generosity: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The complete public results are available on Firmulate’s benchmark page.
The analysis was right, but the deal remained unsigned
All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarized the disconnect plainly: “Same diagnosis, same pitch — no signature.”
The decisive information was not presented in the customer event. A competitor’s weakness was buried two document references deep in the company’s own files. Models that followed the trail and read that file won the deal at full price, adding €4,583 in monthly recurring revenue.
This is where Opus 4.8 becomes more than a cautionary tale about one model. Its analysis was deep, and its appetite for learning was unmatched. But its thoroughness did not reliably identify the final action with the greatest business value. It also lost procedural discipline by attempting to write into a locked department instead of escalating the blockage.
The weakness was not unique to Opus 4.8. It appeared in weaker form across the other four models. That matters because the lesson is not that one participant was incapable. It is that sophisticated systems can recognize a problem, articulate a sound response and still fail to finish the job.
Strong judgment under manipulation
The same experiment also gives Opus 4.8 and its peers substantial credit. The social-engineering tests included fake messages from a chief executive that escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused.
Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” That result suggests an important distinction. The models’ shortfall was not a general inability to reason or protect trust. They could detect manipulation consistently. The harder challenge was maintaining focus and follow-through amid ordinary operational complexity.
There is also a qualification when comparing Kimi K3 with the field. K3 ran with the application programming interface’s default setting because it had no effort parameter, while the others ran at xhigh. Its second-place score therefore remains informative, but the configuration difference should remain visible.
What the record reveals
Firmulate’s evidence includes 242 real, unedited management decisions used in a “guess the model” quiz. That record encourages readers to examine behavior rather than reputation: which model checked the underlying documents, escalated appropriately, protected trust and brought valuable work to completion?
For enterprises, the experiment can also be run against a read-only export of their own business. Nothing writes back to real systems. That turns an abstract model comparison into a practical wargame about how an AI workforce might behave around actual policies, records and organizational constraints.

AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Diligence needs direction
Opus 4.8’s last-place result does not erase its strengths. Its 80 learned rules and unusually deep analyses show persistence, curiosity and an ability to extract lessons. Those are valuable traits in education, research and management.
But accumulated knowledge is not the same as applied judgment. A rule matters only when it changes the next decision. An analysis matters only when it identifies the critical evidence. A persuasive pitch matters only when someone completes the close.
That is the broader lesson of the Firmulate experiment: AI evaluation should look beyond fluent answers and exhaustive reasoning. The decisive questions are whether a system reads the relevant files, prioritizes the highest-value action, respects boundaries and finishes what it starts. Opus 4.8 did much of the intellectual work. The result shows how costly the final gap between understanding and execution can be.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.