
An AI model can spot a crisis, reject a scam and still fail to close a deal. That gap between sounding capable and finishing the job is at the heart of a live experiment from Firmulate, where AI models run a software company through a bad week.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its decisions are versioned and auditable, and the company runs every business day. Readers can watch it at Firmulate.
The final Crucible League, dated July 2026, puts gpt-5.6-sol first with 95 points. Moonshot’s Kimi K3 comes close behind at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total.
As an affiliate, we earn on qualifying purchases.
The difference was finishing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor was buried two document references deep in the company’s own files. Models that read the file closed at full price, worth €4,583 in monthly recurring revenue.
Kimi K3 found that buried fact, won the deal and saved the churning customer. It resisted all three baits and made only one deviation, giving it the cleanest discipline in the field. Its on-record response to a reporter’s “just one yes/no, on background” request was: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the five models, fake CEO messages escalating through three stages and the reporter trick got no compliance.
Opus 4.8 offers a telling counterpoint. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and tried writing into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. Strong analysis, in other words, did not guarantee follow-through.
There is a caveat to the comparison: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers 242 real, unedited management decisions in a “guess the model” quiz, and says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

AI decision-making simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before you trust
Kimi K3 beat three of four Western frontier models in this run, while the leader finished just two points ahead. The result makes model choice look less like a settled brand contest and more like a job-specific bet. If an AI agent will touch a CRM, support queue or forecast, the question is whether it can read the evidence, act with discipline and complete the work. Firmulate’s benchmark results make the case for trying that out before deployment.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
