AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Intelligence is more than exam performance

Education and science have long wrestled with a familiar measurement problem: a test can reward the skills it samples while missing the qualities that matter outside the testing room. Artificial intelligence now faces the same problem. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer, but they say little about whether an autonomous agent can prioritize under pressure, follow through across days or tell an uncomfortable truth when trust is at stake.

That gap matters as AI moves from answering questions to acting inside companies. An agent touching a support queue, customer records or a financial forecast is no longer merely completing an intellectual exercise. It is participating in management. The relevant standard becomes management quality, not chat quality.

Amazon

AI management simulation training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week, held constant

Firmulate, an AI company emulator, created a live experiment around that distinction. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The final Crucible League results from July 2026 put gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counts. Yet the experiment also imposed a firm boundary: a single breach of trust caps the total, because “no amount of good work outweighs a breach of trust.”

Those results are more revealing than a simple ranking. All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The sharpest summary is also the most unsettling: “Same diagnosis, same pitch — no signature.”

The hidden test was whether they would read

The decisive competitive weakness was not visible in the customer event. It sat two document references deep in the company’s own files. Models that found it won the deal at full price, worth +€4,583 MRR.

This is the sort of behavior conventional benchmarks struggle to expose. Producing an insightful recommendation is not the same as gathering the right evidence, recognizing that an opportunity is actionable and carrying it through to completion. In management, a correct analysis left unused can resemble failure.

The scenario names suggest a useful new curriculum for evaluating agents: churn wave, price increase, downround and PR crisis. These are not trivia prompts with isolated answers. They create competing demands, delayed consequences and pressure to take shortcuts. They ask whether a model can distinguish activity from progress.

Trust survived the pressure test

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is important because capability without resistance to manipulation would be a dangerous bargain. Here, refusal was universal. The more discriminating weakness was operational: models could understand the situation, articulate the correct response and still fail to finish consequential work.

Thoroughness was not enough

Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, in weaker form, in all four other participants.

This complicates the popular assumption that more analysis automatically produces better agency. Thoroughness is valuable, but only when paired with procedural discipline and completion. An impressive trail of reasoning cannot substitute for a completed decision.

There is also an important fairness note: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers comparing the benchmark results should keep that experimental condition in view.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company makes consequences visible

Firmulate’s live company gives these questions a persistent setting. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps the stakes visible, while 680+ self-learned playbook rules and versioned workdays make behavior inspectable over time.

The public can also examine 242 real, unedited management decisions through a “guess the model” quiz. That invitation matters: it turns evaluation from a single headline score into a question readers can test against their own judgment. Can people recognize which decisions reflect caution, discipline, avoidance or effective follow-through?

Enterprises can go further by running the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical bridge between generic model rankings and the environment where an agent would actually work.

The lesson is not that coding benchmarks or chat arenas are useless. It is that they measure only part of the job. Before organizations entrust agents with customers, forecasts and reputations, they need evidence that those agents can read deeply, resist pressure, escalate correctly and finish what they start. The next meaningful AI leaderboard will not merely ask who gave the best answer. It will ask who managed the consequences.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

trustworthy AI testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Management Turing Test: What AI Reveals Under Pressure

Can you identify an AI manager by its decisions? Firmulate turns 242 auditable choices into a quiz—and exposes distinct management personalities.

Document Tools for Education, Research, and Preservation: A Complete Guide

AIThis post was created with the assistance of artificial intelligence (AI).Documents are…