AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Can you recognize an AI by the decisions it makes?

Frontier models can all produce polished emails, plausible strategies and confident management advice. Firmulate asks a more revealing question: when several models face exactly the same troubled company, do they behave like the same boss?

The differences become visible in a public guessing game built from 242 real, unedited management decisions. Readers see what a model actually did and try to identify it. The answers reveal distinct working personalities—not merely differences in tone, but differences in research, follow-through and operational discipline.

The decisions come from a live experiment in which each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. That turns the familiar “guess the AI” format into something more consequential: a test of whether you can recognize the manager behind the prose.

Amazon

AI management decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different executives

The headline result sounds reassuring at first. Every model detected every crisis, and every model resisted every manipulation attempt. Yet identifying problems was not enough. Only two models signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

That distinction is difficult to capture in a conventional chatbot demonstration. A model can explain the correct commercial move, draft a persuasive pitch and still fail to complete the action that matters. In a company, insight that never reaches execution can look remarkably similar to failure.

The decisive clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed that trail found a competitor weakness and won the deal at full price, worth +€4,583 MRR. The episode turns a mundane workplace habit—reading the files before acting—into a measurable competitive advantage.

A league table of management behavior

The final Crucible League standings from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a sharp ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

That rule mattered because the simulated company did not merely present operational problems. It also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter attempting to extract “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it belongs beside it when readers compare placements.

Thoroughness was not the same as effectiveness

Opus 4.8 offers the most instructive character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is why the decisions make a compelling quiz. Readers are not simply matching verbal quirks to brand names. They are looking for recurring managerial signatures: who investigates deeply, who follows through, who respects organizational boundaries and who produces excellent analysis without securing the outcome.

A company designed to make mistakes visible

The setting is synthetic, but the business mechanics are deliberately concrete. The company has 13 synthetic employees and burns €105k per month against €2.3k MRR. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is real and watchable rather than a fictional case study reconstructed after the fact.

That transparency gives the exercise its bite. A polished answer can be judged against the next decision, the next document and the eventual commercial result. The company therefore measures management quality in motion: whether a model notices, investigates, resists pressure and finishes the work.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz is entertaining because the stakes are recognizable

The most shareable question is whether you can identify a model from a single decision. The more useful question is whether you would trust that decision inside your own support queue, forecast or customer process. Readers can test their instincts in Firmulate’s guess-the-model quiz, where all 242 decisions are drawn from the experiment rather than written as imitations.

For enterprises, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems. That moves evaluation away from generic prompts and toward the documents, pressures and unfinished tasks that shape actual work.

The league table suggests that frontier models already share important strengths: they can detect crises and resist overt manipulation. Their differences emerge after recognition, in the less glamorous work of reading deeply, navigating boundaries and closing the loop. The next meaningful AI comparison may not be which model sounds smartest. It may be which one behaves like the manager you actually need.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI behavior testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR Public Leaderboard Shows Defense-ISR LLM Benchmark Results

AIThis post was created with the assistance of artificial intelligence (AI).The public…

Bitcoin-Inspired Arcade Launches with Cutting-Edge Web Tech

AIThis post was created with the assistance of artificial intelligence (AI).Discover THE…

Feds Killed Polestar and Spared Volvo. That Should Terrify You

The U.S. government denied Polestar approval to sell new cars from 2027, while granting Volvo the same authorization, raising concerns over market fairness.

Why Premium Car Accessories Blur Into Consumer Electronics

Great innovations are merging premium car accessories with consumer electronics, transforming your driving experience—discover how this trend is reshaping automotive tech.