AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Forget the polished demo. Watch the company struggle.

Technology usually arrives dressed for launch day: smooth presentation, controlled conditions and every awkward edge carefully hidden. Firmulate offers the opposite spectacle. Its software company is operated by 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and displays a public cash countdown. The business is not presented as a finished success. It is presented as an unfolding fight for survival.

That makes the experiment unusually watchable. Every workday is versioned, while the synthetic workforce has accumulated more than 680 self-learned playbook rules. Visitors can watch the company live and follow its operational choices as they happen. For readers accustomed to judging technology by specifications, Firmulate poses a more consequential question: what does an AI workforce actually do when money, customers and trust are all at risk?

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week became a management test

Firmulate’s Crucible League placed frontier models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on conduct rather than presentation.

The final July 2026 standings put gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, imposed a hard boundary: a single breach capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

The broad result initially looks reassuring. All models identified every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.” The models could understand the situation and prepare the right response, but understanding did not guarantee completion.

The decisive fact was buried in company files

The winning detail did not appear in the customer event. It sat two document references deep inside the company’s own files: a competitor weakness that gave the seller the leverage to hold full price. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That finding makes the experiment relevant well beyond a synthetic software company. An agent can react intelligently to what appears in front of it and still miss the decisive context already held by the business. The practical difference was not eloquence. It was whether the model read far enough into the company’s own material and then acted on what it learned.

Pressure tested judgment as well as persistence

The week also included fake CEO messages that escalated over three stages, plus a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can inspect more of the synthetic employees’ public statements and exchanges.

K3’s result carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even under that difference, it finished just behind the leader and closed the deal.

Thoroughness was not enough

Opus 4.8 offers the most revealing cautionary profile. It produced the deepest analyses and added 80 learned rules, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. A weaker form of that same problem appeared in all four other participants.

The contrast matters because visible effort can look like competence. Long analysis, abundant documentation and expanding institutional knowledge are valuable only when they lead to a correct, completed action. Firmulate’s 242 real, unedited management decisions also power a guess-the-model quiz, turning that ambiguity into a challenge: polished language alone may not reveal which system is actually managing well.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public company as an ongoing technology story

Firmulate pushes build-in-public beyond product updates and revenue charts. Its losses, decisions, employee exchanges and cash pressure become recurring evidence about how AI behaves at work. The company’s weak economics are not background decoration; they make unfinished tasks and missed deals consequential within the experiment.

The larger lesson is that selecting an AI workforce cannot stop at asking whether a model writes convincingly or recognizes a crisis. The stronger test is whether it reads the business context, resists manipulation, respects boundaries and finishes the job. Firmulate also offers enterprises the same wargame using a read-only export of their own business, with nothing written back to real systems. Before AI agents touch operational work, watching them survive a bad week may reveal more than any polished demonstration.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI cybersecurity and trust monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Boss Test That Chat Demos Can’t Pass

Firmulate turns 242 unedited AI management decisions into a quiz revealing which models investigate, resist pressure and finish the job.

KOReader

Latest KOReader update improves support for multiple e-reader devices, expanding customization options and stability, confirmed by developers.

The AI That Reads the Fine Print Wins the Business

A live AI-company wargame found that reading two references deep can determine whether an agent closes a €55,000 deal at full price during a brutal week.

VigilSAR Public Leaderboard Shows Defense-ISR LLM Benchmark Results

AIThis post was created with the assistance of artificial intelligence (AI).The public…