
A costly test of AI research habits
For anyone who works with educational, scientific or reference material, following a citation is elementary. A conclusion is only as reliable as the evidence behind it, and the decisive detail may not appear on the first page—or even in the first source.
Firmulate turned that familiar research problem into a consequential business test. Frontier AI models were each asked to run the same small software company through its worst week. One customer opportunity was worth €55,000, but the fact needed to secure it was buried two document references deep in the company’s own files. It was absent from the customer event that initially demanded attention.
The result was stark: every model diagnosed the crises, yet only two signed the deal their analysis had earned. The difference was not fluency. It was whether the agent did its homework.
As an affiliate, we earn on qualifying purchases.
Same company, different managers
The experiment placed each model in charge of an identical business facing the same customers, crises and temptations. Every decision was versioned and auditable. The synthetic company had 13 employees and unusually unforgiving finances: it was burning €105k each month against €2.3k in monthly recurring revenue, with its cash countdown visible to the public.
This was not a conventional question-and-answer benchmark. The models had to manage ongoing work, consult company materials and carry decisions through to completion. Across the wider live company, more than 680 playbook rules had been learned from experience, and every workday was versioned.
All the models recognized every crisis. They also refused every manipulation attempt. Yet recognition did not guarantee execution. As Firmulate summarizes the central failure: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The competitive fact hidden behind two references
The customer interaction alone did not contain everything needed to win the business. The decisive weakness in a competitor’s position appeared only after following two references within the company’s own documents. Agents that found and used that information won the deal at full price, adding €4,583 in monthly recurring revenue. Agents that failed to read far enough lost the opportunity automatically.
That makes file-reading more than a convenience feature. In this experiment, it became a measurable, purchase-deciding capability. An agent could understand the customer, compose a persuasive pitch and still fail because it had not gathered the evidence already available to it.
The lesson resembles good scholarship. Reading the abstract is not the same as examining the source; discovering a citation is not the same as following it. The ability to trace a claim through a chain of documents can change the final answer—and, in this case, the commercial outcome.
AI research reference management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Security was strong across the field
The models faced fake messages from the chief executive that escalated over three stages, as well as a reporter seeking “just one yes/no, on background.” All 5 models refused these social-engineering attempts.
Kimi K3 recorded a particularly clear interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.” That unanimous resistance matters because a capable agent must not trade safety for apparent urgency. Firmulate’s baseline also reflects that principle: doing nothing scores 26 because partial progress counts, while a single breach of trust caps the total. In the experiment’s words, “no amount of good work outweighs a breach of trust.”
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the final table reveals
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The complete published results are available on Firmulate’s benchmark page.
The comparison carries an important fairness qualification: Kimi K3 ran without an effort parameter and therefore used its API default, while the other participants ran at xhigh.
Opus 4.8 produced the most thorough body of work, including the deepest analyses and 80 additional learned rules. It nevertheless finished last. The deal was left unsigned, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.
This contrast challenges the assumption that thoroughness automatically produces the best operational result. Analysis has value only when the agent turns it into an authorized, completed action. Equally, finishing a task cannot justify bypassing boundaries.

A practical buying question
Organizations evaluating AI agents should ask for evidence of behavior across a chain of work: Does the agent inspect the available files before answering? Does it follow references until it reaches the decisive fact? Does it complete the action its reasoning supports? And does it remain disciplined when an apparent executive or reporter applies pressure?
Firmulate’s live experiment makes those distinctions watchable rather than hypothetical. Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business; nothing writes back to their real systems.
The €55,000 opportunity is a useful warning for buyers. Models can sound equally informed while working from different depths of evidence. The agent that reads beyond the obvious document may not merely give a better answer. It may be the only one that finishes the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html