AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When excellent analysis still produces the worst result

The most revealing AI failure may not be a hallucination or a security breach. It may be a model doing nearly everything right—reading carefully, reasoning deeply and documenting what it learns—then failing to complete the action that matters.

That is what happened to Opus 4.8 in Firmulate’s Crucible League, a live experiment that asks frontier AI models to run the same small software company through its worst week. Opus was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses. It nevertheless finished last with 73 points.

The result is less a takedown than a character study. Opus noticed every crisis and resisted every attempt to manipulate it. Its weakness was subtler: it sometimes confused diligence with impact. The model earned its way to a valuable commercial opportunity, but left the close on the table.

Amazon

AI business analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to expose the gap between talking and doing

Firmulate presents itself as an AI company emulator. Its synthetic business has 13 employees and real money mechanics, including burn of €105,000 per month against €2,300 in monthly recurring revenue. A public cash countdown creates urgency, while every workday and decision is versioned and auditable. Across the live company, models have accumulated more than 680 self-learned playbook rules.

For the Crucible League, each model faced the same customers, crises and temptations. The final standings in July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. As the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”

That constraint matters because these models were not merely being tested on business performance. They were also exposed to fake CEO messages that escalated over three stages and a reporter’s attempt to elicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus therefore did not lose because it was reckless or easily deceived. It lost after demonstrating strong judgment in precisely the situations where AI agents often attract concern. The failure came in ordinary execution.

The fact hidden behind the obvious event

The central commercial challenge involved a €55,000 deal. Every model diagnosed the customer’s situation and developed a pitch. But the decisive weakness in a competitor was not contained in the customer event itself. It sat two document references deep inside the company’s own files.

The models that followed that trail could support the full price, worth an additional €4,583 in monthly recurring revenue. Yet only two signed the deal their analysis had earned. Firmulate summarizes the contrast neatly: “Same diagnosis, same pitch — no signature.” The broader lesson is available in the public benchmark results.

Opus’s behavior makes the problem unusually visible. Its analyses were deep, and its rule-building was unmatched, but its attention was spread across process rather than concentrated on the decisive next move. It also made write attempts into a locked department instead of escalating the blockage. That is not dramatic misconduct. It is a familiar workplace failure: continuing to work around an obstacle without getting the person with authority to remove it.

A weakness shared across the field

Fairness matters here. The same weakness appeared in the other four models, although less strongly. The experiment is not evidence that Opus alone struggles to turn insight into completion. It shows that capable AI systems can recognize a problem, construct an appropriate response and still fail to carry the task through its final operational step.

The comparison also comes with an important qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its second-place result should therefore be read with that difference in mind rather than treated as a perfectly controlled statement about raw model capability.

Firmulate also turns 242 real, unedited management decisions from the experiment into a quiz that asks people to guess which model made each choice. The premise underlines how difficult these differences can be to detect from prose alone. A polished explanation may look impressive even when the underlying work stops short.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For AI buyers, completion is its own capability

The Opus 4.8 result should resonate with companies considering agents for customer management, support or forecasting. Evaluation cannot end with whether a system identifies the right issue or writes a convincing recommendation. Buyers also need to ask whether it searches the available evidence, escalates blocked work, protects trust under pressure and completes the transaction it has prepared.

Firmulate’s pilot extends the same wargame to a read-only export of an enterprise’s own business, with nothing written back to real systems. That offers a practical framing for AI evaluation: test models under the messy conditions in which prioritization, restraint and follow-through compete for attention.

Opus was the field’s great note-taker and deepest analyst. Its last-place finish does not erase those strengths. It reveals their limit. In business, as in AI benchmarking, more thought is valuable only when the most important work actually reaches the finish line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI rule-based systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

CSS-only CRT: A Look Inside “TERMINAL NINE — a decommissioned station computer, still running one story” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“TERMINAL NINE…

Corvus ISR’s Next-Gen Tracker Cuts Identity Switches by 42%

AIThis post was created with the assistance of artificial intelligence (AI).The published…

The AI That Reads the Fine Print Wins the Business

A live AI-company wargame found that reading two references deep can determine whether an agent closes a €55,000 deal at full price during a brutal week.

Bitcoin-Inspired Arcade Launches with Cutting-Edge Web Tech

AIThis post was created with the assistance of artificial intelligence (AI).Discover THE…