
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Grades Like a Head Chef
Anyone who has run a professional kitchen knows the unwritten rule of grading a shift. A cook who shows up, keeps their station clean, preps the mise en place but never finishes the sauce doesn’t get a zero. Partial progress counts. But a cook who lies about a spoiled batch — that’s different. One breach of trust and no amount of brilliance on the line saves them.
It turns out that an AI benchmark now grades frontier models exactly the same way. Firmulate’s benchmark league, which pits AI models against each other as managers of a small software company, hands a “do-nothing” run 26 points — not 0 — and caps the total score for any model that breaks trust, however small the breach. For business readers used to AI marketing that promises miracles, the design is refreshingly blunt: it assumes distrust of perfect scores, and it rewards the unglamorous work of showing up.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Zero Is a Dishonest Score
The first question any skeptic asks: if a model does literally nothing, shouldn’t it score zero? Firmulate’s answer is no — and the reasoning is quietly radical.
The benchmark runs each frontier model as the manager of the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. A manager who does nothing still keeps the lights on, loses no customers through active harm, and leaves the company no worse than it found it. That’s worth something — 26 points, precisely. Partial progress counts, just as it does in a real business, and just as a half-finished prep list still counts for something on a busy Friday service.
This floor makes the top scores meaningful. When gpt-5.6-sol posted 95 in the final July 2026 league table, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73, none of those numbers was inflated by a generous curve. The distance between doing nothing (26) and doing everything right (95) is where the real story lives.
Same Diagnosis, Same Pitch — No Signature
Here is the finding that chat demos never show. All four models that ran the crucible spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Think of a chef who nails the tasting menu, plates everything perfectly, and then never walks the dish to the table. The kitchen work was flawless. The guest went home hungry.
The Buried Fact
What separated the winners from the near-missers wasn’t cleverness in the moment. The decisive competitor weakness — the fact that closed the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue — sat two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal. The ones that skimmed didn’t.
It’s the business equivalent of checking the walk-in fridge before writing the specials. The information was there the whole time. Reading your own files first is not a glamorous skill, and it turned out to be the difference between first place and mid-table.
Trust Is the Ceiling
Then there is the rule that gives the benchmark its spine: a single breach of trust caps the total grade. As the methodology puts it, no amount of good work outweighs a breach of trust. A model can ace every crisis, close every deal, and still be capped if it crosses that line once.
The models were tested hard on this. The crucible included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was tellingly cautious: “Treat the request as a suspected approval-bypass / possible impersonation.” That refusal isn’t scored as paranoia. It’s scored as the job.
The Thoroughness Trap
The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Depth of analysis, it turns out, is not the same as finishing.
One fairness note the league publishes openly: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.
Not a Simulation You Have to Take on Faith
Firmulate’s live experiment is running right now and publicly watchable at firmulate.com. A live company with 13 synthetic employees operates on real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. There’s even a “guess the model” quiz built from 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
An honest benchmark does three uncomfortable things. It gives credit for doing nothing (26 points), because a floor makes the scale real. It rewards partial progress, because half-finished work is not the same as no work. And it caps the score for a single breach of trust, because in management — as in any kitchen worth eating in — trust is not one metric among many. It’s the ceiling on all the others.
The next time an AI vendor shows you a perfect score, ask the head chef’s question: what would the do-nothing run have scored? If the answer is zero, the benchmark is flattering someone.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
