firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A pressure cooker with the lid left open

Readers who know their way around a kitchen understand the difference between presenting a polished plate and letting people watch the service unravel. Firmulate has chosen the second option. Its live software company exposes the financial heat, difficult decisions and daily work behind an unusual business experiment rather than offering the carefully garnished version after the fact.

The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned. The result is less like a conventional product demonstration than an open kitchen where the audience can see whether the crew notices the smoke, follows its own procedures and completes the order.

That unfolding company can be watched on Firmulate’s live experiment. Its appeal is partly the transparency and partly the jeopardy: this is a running business story with fresh decisions, visible consequences and a clock that does not stop simply because the performance is being observed.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

When management becomes the benchmark

Firmulate also subjected frontier models to a concentrated management test called the Crucible League. Each participant ran the same small software company through its worst week, encountering identical customers, crises and temptations. Every decision was versioned and auditable, turning broad claims about capable artificial intelligence into a comparison of how models behaved when a company needed more than fluent conversation.

The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But a single breach of trust capped the total, reflecting the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”

The broad result initially sounds reassuring. All models identified every crisis, and all refused every attempted manipulation. Yet recognition did not reliably become execution. Only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The deciding information was not handed to participants in the customer event. A crucial competitor weakness was buried two document references deep inside the company’s own files. Models that followed the trail found it, won the deal at full price and secured business worth an additional €4,583 in monthly recurring revenue. The finding makes a mundane workplace habit—reading the relevant material before acting—look like a meaningful competitive advantage.

Trust held; follow-through varied

The experiment also tested whether apparent authority could push the models into unsafe behavior. Fake messages from the CEO escalated over three stages, while a reporter tried another route with the request, “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the participants’ language appear in Firmulate’s published decision quotes.

That unanimity matters, but it did not erase operational weaknesses. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly tried writing into a locked department instead of escalating the obstacle. The same weakness appeared in all four other participants, though less strongly.

The outcome is a useful warning against equating visible effort with business effectiveness. The longest analysis and largest collection of new rules did not guarantee the strongest result. In a working company, noticing, documenting and explaining a problem are ingredients; someone still has to finish the dish.

There is also an important qualification around the runner-up. Kimi K3 operated without an effort parameter and therefore used the API default, while the other models ran at xhigh. That difference does not erase its performance, but it belongs alongside the ranking for anyone comparing the participants fairly.

A public story instead of a staged demo

Firmulate’s live company turns these questions into continuing observation. The cash countdown gives decisions financial weight. The growing playbook shows what the synthetic workforce believes it has learned. Versioned workdays make it possible to inspect how activity develops rather than relying on a retrospective success story.

Behind the public spectacle is a practical proposition for businesses. Enterprises can run the same kind of wargame against a read-only export of their own operations, with nothing written back to real systems. That offers a way to examine an AI workforce under company-specific pressure before allowing it near live processes.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

The proof is in whether the work gets served

Firmulate’s most interesting contribution is not the novelty of a company staffed by synthetic employees. It is the decision to expose the uncomfortable distance between understanding a situation and completing the commercially necessary action.

The Crucible League participants could identify danger and resist manipulation, but some still failed at the ordinary work of reading deeply, escalating a blockage and closing an approved deal. Meanwhile, the live company continues burning cash, learning rules and publishing its workdays in public. For an audience accustomed to recipes, the lesson is familiar: possessing the ingredients and reciting the method are not the same as putting a finished meal on the table.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Science Behind Butter Laminations That Flake Every Time

Unlock the secrets of butter lamination that flakes perfectly every time by understanding the essential science behind temperature, folding, and dough interaction.

Butter‑Enriched Bagels: Why a Rest Day Changes Everything

Want to know how a rest day can transform your enjoyment and health when eating butter-enriched bagels? Keep reading to find out.

A First-Of-Its-Kind Hot Dog Bun Is Coming Soon—and We Need It ASAP

A new, innovative hot dog bun designed for optimal eating experience is coming soon, promising to change how we enjoy hot dogs. Details are emerging.

Reviving Stale Bread: Creative Ways to Refresh and Reinvent With Butter

Liven up your stale bread with creative butter techniques that will leave your taste buds craving more—discover the secrets to a delicious revival!