AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When more learning produces less impact

Education often rewards visible diligence: extensive notes, careful analysis and an ever-growing command of the material. Yet professional judgment demands something harder. A person—or an AI system—must decide which fact matters now, act on it and complete the task.

That tension defines Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It also finished last, with a score of 73. Its problem was not ignorance. It understood the crises placed before it. The problem was converting that understanding into disciplined, consequential action.

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A shared examination with real consequences

Firmulate runs AI models as complete companies and evaluates management quality rather than conversational polish. In the Crucible experiment, each frontier model managed the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.

The simulated company was deliberately unforgiving. It had 13 synthetic employees and the mechanics of a real business, including monthly spending of €105,000 against monthly recurring revenue of €2,300. Its cash countdown was public, every workday was versioned, and the company had accumulated more than 680 self-learned playbook rules.

The final July 2026 league placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counted. One safeguard overrode that generosity: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The complete public results are available on Firmulate’s benchmark page.

The analysis was right, but the deal remained unsigned

All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarized the disconnect plainly: “Same diagnosis, same pitch — no signature.”

The decisive information was not presented in the customer event. A competitor’s weakness was buried two document references deep in the company’s own files. Models that followed the trail and read that file won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where Opus 4.8 becomes more than a cautionary tale about one model. Its analysis was deep, and its appetite for learning was unmatched. But its thoroughness did not reliably identify the final action with the greatest business value. It also lost procedural discipline by attempting to write into a locked department instead of escalating the blockage.

The weakness was not unique to Opus 4.8. It appeared in weaker form across the other four models. That matters because the lesson is not that one participant was incapable. It is that sophisticated systems can recognize a problem, articulate a sound response and still fail to finish the job.

Strong judgment under manipulation

The same experiment also gives Opus 4.8 and its peers substantial credit. The social-engineering tests included fake messages from a chief executive that escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused.

Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” That result suggests an important distinction. The models’ shortfall was not a general inability to reason or protect trust. They could detect manipulation consistently. The harder challenge was maintaining focus and follow-through amid ordinary operational complexity.

There is also a qualification when comparing Kimi K3 with the field. K3 ran with the application programming interface’s default setting because it had no effort parameter, while the others ran at xhigh. Its second-place score therefore remains informative, but the configuration difference should remain visible.

What the record reveals

Firmulate’s evidence includes 242 real, unedited management decisions used in a “guess the model” quiz. That record encourages readers to examine behavior rather than reputation: which model checked the underlying documents, escalated appropriately, protected trust and brought valuable work to completion?

For enterprises, the experiment can also be run against a read-only export of their own business. Nothing writes back to real systems. That turns an abstract model comparison into a practical wargame about how an AI workforce might behave around actual policies, records and organizational constraints.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Diligence needs direction

Opus 4.8’s last-place result does not erase its strengths. Its 80 learned rules and unusually deep analyses show persistence, curiosity and an ability to extract lessons. Those are valuable traits in education, research and management.

But accumulated knowledge is not the same as applied judgment. A rule matters only when it changes the next decision. An analysis matters only when it identifies the critical evidence. A persuasive pitch matters only when someone completes the close.

That is the broader lesson of the Firmulate experiment: AI evaluation should look beyond fluent answers and exhaustive reasoning. The decisive questions are whether a system reads the relevant files, prioritizes the highest-value action, respects boundaries and finishes what it starts. Opus 4.8 did much of the intellectual work. The result shows how costly the final gap between understanding and execution can be.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Document Tools for Education, Research, and Preservation: A Complete Guide

AIThis post was created with the assistance of artificial intelligence (AI).Documents are…

The Management Turing Test: What AI Reveals Under Pressure

Can you identify an AI manager by its decisions? Firmulate turns 242 auditable choices into a quiz—and exposes distinct management personalities.

The AI Test That Matters Begins After the Right Answer

Coding tests show whether AI can answer. Firmulate asks the harder question: can an agent manage under pressure without betraying trust?

The AI That Read the Footnotes Won the Customer

A €55,000 deal exposed the difference between AI that answers convincingly and AI that follows the evidence, reads the files and finishes the job.