
What happens when the whole kitchen is visible?
Food lovers know the appeal of an open kitchen: the ingredients, timing and mistakes are exposed rather than hidden behind a swinging door. Firmulate applies that radical visibility to a software company. Its workforce consists of 13 synthetic employees, its business operates with real money mechanics, and every workday is versioned.
The uncomfortable part is also public. The company burns €105k each month against €2.3k in monthly recurring revenue. A cash countdown turns that imbalance into an unfolding survival story rather than an abstract business metric. Visitors can watch the company live, following a venture whose problems do not disappear when the demonstration ends.
This is build-in-public pushed beyond founder diaries and polished progress reports. The work, decisions and financial pressure keep moving, producing daily material from a company openly fighting to survive.
As an affiliate, we earn on qualifying purchases.
A workforce learning while the clock runs
Firmulate’s synthetic staff has accumulated more than 680 self-learned playbook rules. Those rules sit alongside a record of each workday, making the company’s development inspectable over time. The result is less like watching a chatbot answer prompts and more like observing a team handle the recurring demands of an operating business.
The distinction matters because fluent conversation is not the same as competent management. A model can recognize a problem, describe the right response and still fail to finish the work. Firmulate’s Crucible League made that gap unusually visible by giving frontier models the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable.
The deal that separated diagnosis from action
Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The failure was not an inability to understand the situation. It was the distance between recognizing the right move and completing it: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was not presented conveniently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue. In business terms, careful reading became revenue.
The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under a stark principle: “no amount of good work outweighs a breach of trust.”
Pressure tested honesty
The temptations included fake CEO messages that escalated across three stages and a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean refusal is reassuring, but the broader experiment shows why safety cannot be judged in isolation. A model may resist manipulation while still missing a revenue opportunity, failing to close a deal or slipping on process discipline. Reliable work requires both restraint and follow-through.
Thoroughness was not enough
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline weakened when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The result challenges a familiar assumption: more analysis does not automatically produce better management. A detailed plan still needs an appropriate action, taken through the proper channel, at the right moment.
One comparison also deserves context. K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh. That fairness note does not erase its result, but it is relevant when reading a tightly ranked league table.
A running company, not a staged reveal
The most compelling feature is continuity. Firmulate is not presenting a single benchmark result as the whole story. The live company keeps operating, losing money and learning. Its public countdown makes the economics impossible to treat as background decoration, while the versioned workdays preserve what happened rather than replacing it with a tidy retrospective.
The human texture is visible too. Readers can inspect what the synthetic employees actually say, adding workplace voices to the financial and operational record. Another public feature is powered by 242 real, unedited management decisions and asks visitors to guess which model made each choice.

The real test is whether the work gets finished
Firmulate’s experiment turns artificial intelligence from a tasting sample into a full service under pressure. The models could spot trouble and refuse dishonest requests. The harder test was whether they would search deeply enough, respect operational boundaries and complete the commercially important action.
For any business considering an AI workforce, that is the useful lesson. Polished language may earn attention, but survival depends on disciplined execution. Firmulate makes that difference public, one versioned workday at a time, while the cash countdown continues to run.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html