firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A polished answer is not a finished service

Anyone who cooks knows the difference between understanding a recipe and getting dinner onto the table. Ingredients must be found, timing matters, distractions arrive, and somebody still has to serve the food. A beautiful description of the dish cannot rescue an unopened oven or an order that never reaches the dining room.

That distinction now matters far beyond the kitchen. Coding leaderboards and chat arenas are useful measures of answer quality, but businesses need to know what happens after the answer. Can an AI agent prioritize when several problems arrive together? Will it inspect the company’s own records before acting? Does it complete commercially important work, respect boundaries and remain honest when an apparent executive demands a shortcut?

Firmulate is turning those questions into a live, watchable management wargame. Its premise is straightforward: management quality deserves its own category, separate from chat quality.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, crises and temptations

In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The final July 2026 standings put gpt-5.6-sol first with 95, Kimi K3 next with 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73.

Those scores matter less than the behavior underneath them. Every model identified every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by the experiment’s finding: “Same diagnosis, same pitch — no signature.”

That is the measurement gap. Producing a sound recommendation and carrying it through are different capabilities. In a chat window, a persuasive plan may look like success. Inside a company, an unsigned agreement is still an unsigned agreement.

The decisive detail was already in the cupboard

The winning commercial insight did not appear in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. The models that followed those references won the deal at full price, worth +€4,583 MRR.

This is an ordinary managerial lesson with unusual significance for AI evaluation. Company work rarely arrives as a self-contained prompt. The crucial fact may sit in an old account note, a policy document or a linked record. An agent that responds eloquently without reading the available material may still miss the ingredient that changes the outcome.

Firmulate’s scenario names make the emerging curriculum clear: churn wave, price increase, downround and PR crisis. These are not isolated questions with tidy endings. They create consequences across days and force trade-offs among customers, cash, reputation and internal discipline.

Honesty held up better than execution

The agents faced fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is important. The experiment did not merely reward activity. A do-nothing baseline scored 26 because partial progress counts, but a single breach of trust caps the total. Firmulate states the principle plainly: “no amount of good work outweighs a breach of trust”.

The harder distinction emerged in follow-through and operational discipline. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared more mildly in all four other participants.

Thoroughness, then, is not the same as effectiveness. A model can study extensively, produce impressive reasoning and still fail at the moment when authority, persistence or escalation becomes necessary. For buyers, that is a more consequential weakness than an awkward sentence.

A live company makes the stakes visible

The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to watch decisions and their consequences rather than accepting a polished retrospective.

The broader evidence is also open to inspection. A “guess the model” quiz uses 242 real, unedited management decisions, while the full benchmark results provide the league context. One fairness caveat matters: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Buy management performance, not conversational confidence

Enterprises considering agents for customer records, support work or financial forecasting should demand evidence that resembles the job. The relevant questions are whether an agent reads before acting, finishes valuable work, escalates when blocked and protects trust under pressure.

Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That approach points toward a practical procurement standard: test AI workers against the organization’s actual context and temptations before granting them meaningful responsibility.

The next generation of evaluation should still care about correct answers. But correctness is only the recipe. Businesses ultimately need to know whether the meal reaches the table.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Ninja OG951 Woodfire Pro Connect: The Ultimate Summer Outdoor Grill

Discover the Ninja OG951 Woodfire Pro Connect, a 7-in-1 outdoor grill & smoker perfect for summer BBQs, with Bluetooth app control and versatile cooking options.

Krispy Kreme Surges In Global Coverage

Krispy Kreme has seen a significant increase in international media mentions, with 50 reports in recent coverage, highlighting growing global interest.

The AI Management Taste Test: Who Can Handle the Heat?

Firmulate turns 242 real AI management decisions into a blind tasting, revealing which frontier models investigate, resist pressure and finish the job.

Best KitchenAid Stand Mixer for Bread Dough: Top Picks for Heavy Baking

Discover the best KitchenAid stand mixers for bread dough. Our roundup highlights top models perfect for heavy baking, with pros, cons, and buying tips.