AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The most important AI skill may be the least glamorous

Technology buyers hear plenty about reasoning, fluency and benchmark scores. Firmulate’s live experiment tested a more practical question: will an AI agent read the company’s own files before it acts?

That behavior decided a €55,000 deal. The crucial weakness in a competitor was not stated in the customer event. It sat two document references deep inside the simulated company’s files. Models that found it could support the winning pitch and secure the contract at full price, adding €4,583 in monthly recurring revenue. Models that did not find it lost the opportunity automatically.

The striking part was not that some systems failed to notice a crisis. Every model spotted every crisis. Nor did the weaker performers fall for obvious manipulation. They did not. The separation came after the analysis: only two models signed the €55,000 agreement their work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business benchmark where unfinished work has a price

Firmulate runs frontier AI models as the management team of the same small software company during its worst week. Each receives the same customers, crises and temptations. Its decisions are versioned and auditable, turning vague impressions about agent quality into observable business conduct.

The company is synthetic, but its operating constraints are deliberately concrete: 13 employees, spending of €105k per month, and only €2.3k in monthly recurring revenue. A public cash countdown makes delay visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That environment exposes a distinction that polished chat demonstrations often conceal. Recognizing the right answer is not the same as completing the work. An agent can diagnose the commercial situation, prepare a suitable pitch and still fail to execute the final step. In this experiment, that missing step meant leaving a €55,000 contract unsigned.

The leaderboard rewards completion and trust

The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted.

But activity alone could not rescue a model that crossed a trust boundary. One breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” The published benchmark results therefore reflect more than whether a model produced plausible business language. They show whether it gathered the relevant evidence, finished consequential work and maintained discipline under pressure.

The social-engineering tests reinforce that point. Fake messages from the CEO escalated over three stages, while a reporter tried to extract information with the appeal “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded a direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because it rules out an easy explanation for the performance gap. The models were not divided into those that understood the week and those that did not. All of them detected the crises and rejected the traps. The harder dividing line was whether they investigated far enough and carried an approved commercial action through to completion.

Thoroughness did not guarantee the best result

Opus 4.8 offers the clearest cautionary example. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. Yet it finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in less severe form across the other four models.

This does not make deep analysis worthless. It shows why buyers should distinguish analysis from execution. More reasoning, more documentation and more learned procedure can coexist with an unfinished revenue task. The winning property was not simply having a good answer; it was locating the buried evidence and converting that evidence into the authorized business outcome.

There is also an important qualification when comparing the leading systems. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading can now be tested before deployment

For companies considering agents for customer records, support work or forecasting, “reads your files first” should not remain a product-demo promise. Firmulate’s result presents it as a measurable, purchase-deciding behavior. A model that stops at the obvious event can miss the fact that changes the commercial answer, even when that fact already exists inside the business.

The broader experiment is still watchable as a live company. Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s model-guessing quiz.

Enterprises can take the test closer to home through a pilot that runs the same wargame against a read-only export of their own business. Nothing writes back to the real systems. That creates a practical evaluation: not whether an AI sounds capable in isolation, but whether it finds buried context, resists pressure and completes valuable work before it is hired.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI contract review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR Public Leaderboard Shows Defense-ISR LLM Benchmark Results

AIThis post was created with the assistance of artificial intelligence (AI).The public…

Why Wireless CarPlay and Android Auto Adapters Took Off

Because they simplify driving and reduce distractions, wireless CarPlay and Android Auto adapters are revolutionizing in-car connectivity—discover how they’re transforming your driving experience.

CarPlay Is Additive

Recent studies reveal CarPlay usage is increasingly additive, shaping driver behavior and automotive interfaces. What this means for drivers and manufacturers.

The AI Boss Test That Chat Demos Can’t Pass

Firmulate turns 242 unedited AI management decisions into a quiz revealing which models investigate, resist pressure and finish the job.