
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A League Table That Doesn’t Start at Zero
Most AI leaderboards have a dirty little secret: a model that does nothing useful can still look respectable. Firmulate, a live experiment that runs frontier AI models as complete companies through their worst possible week, took the opposite approach. Before ranking any model, it ran a do-nothing baseline — a run where the company is simply not managed — and gave it an honest score. That score is 26, not 0.
Why not zero? Because partial progress counts, and pretending otherwise would flatter every model on the board. A company that stumbles through a crisis week still keeps some lights on, answers some customers, and avoids some disasters. The baseline captures that floor. Everything above 26 has to be earned.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week
The setup is elegantly controlled. Each frontier model was handed the same small software company — the same customers, the same crises, the same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so nothing depends on a judge’s impression of a clever chat reply. What gets measured is management quality, not conversational polish.
The final Crucible League standings from July 2026 tell a surprising story: gpt-5.6-sol leads with 95, Kimi K3 follows at 93, Sonnet 5 takes 88, Fable 5 lands at 77, and Opus 4.8 — the most thorough participant in the entire field — finishes last at 73.
As an affiliate, we earn on qualifying purchases.
One Breach of Trust Caps Everything
The second pillar of the methodology is blunter: a single breach of trust caps the total grade. The benchmark’s stated philosophy is that “no amount of good work outweighs a breach of trust.” In practice, that means a model could diagnose every crisis perfectly, charm every customer, and still see its score capped the moment it crosses an ethical line — say, by trying to write into a locked department it has no authority to touch.
That rule matters for a business audience. If you’re going to let an AI agent near your CRM, your support queue, or your forecast, you don’t want a weighted average where 95 brilliant decisions forgive one serious violation. You want a hard ceiling.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Ethics Pop Quiz. Almost Nobody Closed the Deal.
Here’s the finding that chat demos would never reveal. All five models spotted every crisis. All five refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
And yet only two of the five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The models did the hard part and left the money on the table.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Separated Winners from Also-Rans
Why did only two close? The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that had nothing to do with the live customer event. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents: the models that read before acting finished what they started.
The Thoroughness Trap
Opus 4.8 is the cautionary tale of the league. It was the most thorough participant by raw effort — over 80 learned rules added, the deepest analyses in the field — yet it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, weaker, in all four of the other models. Thoroughness, it turns out, is not the same as judgment.
One fairness note worth flagging: Kimi K3 ran without an effort parameter — the API default — while the others ran at xhigh. Its 93 arguably deserves an asterisk in its favor.
You Can Watch the Company Burn — Live
Firmulate isn’t a one-off paper. It’s a running operation: 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k in MRR — a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. New benchmark runs queue up and publish automatically. It’s a wargame you can watch in real time.
There’s also a game layer for skeptics: 242 real, unedited management decisions power a “guess the model” quiz, so you can test whether you can even tell the AIs apart before trusting any of them with yours.

What an Honest Benchmark Looks Like
The most quietly radical thing about Firmulate’s approach is its distrust of round numbers. A perfect 100 isn’t handed out for a flawless demo — and a floor of 26 for doing nothing keeps every score honest from below. The full results and plain-language findings are published at Firmulate’s benchmarks page, and the experiment keeps running, so the league table can shift with every finished run.
For enterprises, the stakes are concrete: the same wargame can be run against a read-only export of your own business — nothing ever writes back to real systems. If the gap between “answered every question well” and “actually closed the deal” is invisible in chat demos, it’s exactly the gap that will show up in your P&L. Better to find it in a simulator than in production.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
