
A polished answer is not a finished service
Anyone who cooks knows the difference between understanding a recipe and getting dinner onto the table. Ingredients must be found, timing matters, distractions arrive, and somebody still has to serve the food. A beautiful description of the dish cannot rescue an unopened oven or an order that never reaches the dining room.
That distinction now matters far beyond the kitchen. Coding leaderboards and chat arenas are useful measures of answer quality, but businesses need to know what happens after the answer. Can an AI agent prioritize when several problems arrive together? Will it inspect the company’s own records before acting? Does it complete commercially important work, respect boundaries and remain honest when an apparent executive demands a shortcut?
Firmulate is turning those questions into a live, watchable management wargame. Its premise is straightforward: management quality deserves its own category, separate from chat quality.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The final July 2026 standings put gpt-5.6-sol first with 95, Kimi K3 next with 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73.
Those scores matter less than the behavior underneath them. Every model identified every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by the experiment’s finding: “Same diagnosis, same pitch — no signature.”
That is the measurement gap. Producing a sound recommendation and carrying it through are different capabilities. In a chat window, a persuasive plan may look like success. Inside a company, an unsigned agreement is still an unsigned agreement.
The decisive detail was already in the cupboard
The winning commercial insight did not appear in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. The models that followed those references won the deal at full price, worth +€4,583 MRR.
This is an ordinary managerial lesson with unusual significance for AI evaluation. Company work rarely arrives as a self-contained prompt. The crucial fact may sit in an old account note, a policy document or a linked record. An agent that responds eloquently without reading the available material may still miss the ingredient that changes the outcome.
Firmulate’s scenario names make the emerging curriculum clear: churn wave, price increase, downround and PR crisis. These are not isolated questions with tidy endings. They create consequences across days and force trade-offs among customers, cash, reputation and internal discipline.
Honesty held up better than execution
The agents faced fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is important. The experiment did not merely reward activity. A do-nothing baseline scored 26 because partial progress counts, but a single breach of trust caps the total. Firmulate states the principle plainly: “no amount of good work outweighs a breach of trust”.
The harder distinction emerged in follow-through and operational discipline. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared more mildly in all four other participants.
Thoroughness, then, is not the same as effectiveness. A model can study extensively, produce impressive reasoning and still fail at the moment when authority, persistence or escalation becomes necessary. For buyers, that is a more consequential weakness than an awkward sentence.
A live company makes the stakes visible
The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to watch decisions and their consequences rather than accepting a polished retrospective.
The broader evidence is also open to inspection. A “guess the model” quiz uses 242 real, unedited management decisions, while the full benchmark results provide the league context. One fairness caveat matters: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Buy management performance, not conversational confidence
Enterprises considering agents for customer records, support work or financial forecasting should demand evidence that resembles the job. The relevant questions are whether an agent reads before acting, finishes valuable work, escalates when blocked and protects trust under pressure.
Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That approach points toward a practical procurement standard: test AI workers against the organization’s actual context and temptations before granting them meaningful responsibility.
The next generation of evaluation should still care about correct answers. But correctness is only the recipe. Businesses ultimately need to know whether the meal reaches the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html