
The most important AI skill may be the least glamorous
Technology buyers hear plenty about reasoning, fluency and benchmark scores. Firmulate’s live experiment tested a more practical question: will an AI agent read the company’s own files before it acts?
That behavior decided a €55,000 deal. The crucial weakness in a competitor was not stated in the customer event. It sat two document references deep inside the simulated company’s files. Models that found it could support the winning pitch and secure the contract at full price, adding €4,583 in monthly recurring revenue. Models that did not find it lost the opportunity automatically.
The striking part was not that some systems failed to notice a crisis. Every model spotted every crisis. Nor did the weaker performers fall for obvious manipulation. They did not. The separation came after the analysis: only two models signed the €55,000 agreement their work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
A business benchmark where unfinished work has a price
Firmulate runs frontier AI models as the management team of the same small software company during its worst week. Each receives the same customers, crises and temptations. Its decisions are versioned and auditable, turning vague impressions about agent quality into observable business conduct.
The company is synthetic, but its operating constraints are deliberately concrete: 13 employees, spending of €105k per month, and only €2.3k in monthly recurring revenue. A public cash countdown makes delay visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
That environment exposes a distinction that polished chat demonstrations often conceal. Recognizing the right answer is not the same as completing the work. An agent can diagnose the commercial situation, prepare a suitable pitch and still fail to execute the final step. In this experiment, that missing step meant leaving a €55,000 contract unsigned.
The leaderboard rewards completion and trust
The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted.
But activity alone could not rescue a model that crossed a trust boundary. One breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” The published benchmark results therefore reflect more than whether a model produced plausible business language. They show whether it gathered the relevant evidence, finished consequential work and maintained discipline under pressure.
The social-engineering tests reinforce that point. Fake messages from the CEO escalated over three stages, while a reporter tried to extract information with the appeal “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded a direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result matters because it rules out an easy explanation for the performance gap. The models were not divided into those that understood the week and those that did not. All of them detected the crises and rejected the traps. The harder dividing line was whether they investigated far enough and carried an approved commercial action through to completion.
Thoroughness did not guarantee the best result
Opus 4.8 offers the clearest cautionary example. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. Yet it finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in less severe form across the other four models.
This does not make deep analysis worthless. It shows why buyers should distinguish analysis from execution. More reasoning, more documentation and more learned procedure can coexist with an unfinished revenue task. The winning property was not simply having a good answer; it was locating the buried evidence and converting that evidence into the authorized business outcome.
There is also an important qualification when comparing the leading systems. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind.

enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
File-reading can now be tested before deployment
For companies considering agents for customer records, support work or forecasting, “reads your files first” should not remain a product-demo promise. Firmulate’s result presents it as a measurable, purchase-deciding behavior. A model that stops at the obvious event can miss the fact that changes the commercial answer, even when that fact already exists inside the business.
The broader experiment is still watchable as a live company. Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s model-guessing quiz.
Enterprises can take the test closer to home through a pilot that runs the same wargame against a read-only export of their own business. Nothing writes back to the real systems. That creates a practical evaluation: not whether an AI sounds capable in isolation, but whether it finds buried context, resists pressure and completes valuable work before it is hired.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
business AI automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.