AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

An AI model can spot a crisis, reject a scam and still fail to close a deal. That gap between sounding capable and finishing the job is at the heart of a live experiment from Firmulate, where AI models run a software company through a bad week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its decisions are versioned and auditable, and the company runs every business day. Readers can watch it at Firmulate.

The final Crucible League, dated July 2026, puts gpt-5.6-sol first with 95 points. Moonshot’s Kimi K3 comes close behind at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total.

Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference was finishing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive weakness in a competitor was buried two document references deep in the company’s own files. Models that read the file closed at full price, worth €4,583 in monthly recurring revenue.

Kimi K3 found that buried fact, won the deal and saved the churning customer. It resisted all three baits and made only one deviation, giving it the cleanest discipline in the field. Its on-record response to a reporter’s “just one yes/no, on background” request was: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the five models, fake CEO messages escalating through three stages and the reporter trick got no compliance.

Opus 4.8 offers a telling counterpoint. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and tried writing into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. Strong analysis, in other words, did not guarantee follow-through.

There is a caveat to the comparison: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers 242 real, unedited management decisions in a “guess the model” quiz, and says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before you trust

Kimi K3 beat three of four Western frontier models in this run, while the leader finished just two points ahead. The result makes model choice look less like a settled brand contest and more like a job-specific bet. If an AI agent will touch a CRM, support queue or forecast, the question is whether it can read the evidence, act with discipline and complete the work. Firmulate’s benchmark results make the case for trying that out before deployment.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deployment testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR Public Leaderboard Shows Defense-ISR LLM Benchmark Results

AIThis post was created with the assistance of artificial intelligence (AI).The public…

The Fake Boss Came Calling. The AI Agents Didn’t Flinch

Five frontier AI models faced escalating fake-CEO demands and a reporter’s trick. All refused, showing integrity can be tested before deployment.

CSS-only CRT: A Look Inside “TERMINAL NINE — a decommissioned station computer, still running one story” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“TERMINAL NINE…

Why Automotive Convenience Became a Gadget Story

Why automotive convenience evolved into a gadget story, revealing how smart technology transforms driving into a personalized, safer experience—discover the future of mobility.