firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A recipe can survive a missing pinch of salt. A restaurant, like any small business, may not survive a missed warning, a lost customer or a deal left unsigned. Firmulate’s live experiment puts AI models in that kind of pressure cooker: run the same software company through its worst week and see whether they finish the job.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The newcomer nearly tops the table

In Firmulate’s final Crucible League results for July 2026, Moonshot’s Kimi K3 placed second with 93 points, two behind gpt-5.6-sol at 95. It beat Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. That is a striking result for a newcomer: K3 finished ahead of three of the four Western frontier models in the comparison.

The experiment gave each model the same small software company, the same customers, crises and temptations. Its decisions were versioned and auditable. The company itself has 13 synthetic employees and real money mechanics: it is burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. More than 680 self-learned playbook rules and every workday’s activity are part of a live operation at Firmulate.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Seeing the problem wasn’t the hard part

All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The deal depended on a buried weakness in a competitor’s offer, two document references deep in the company’s own files rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue.

That gap between diagnosis and action is the experiment’s sharpest finding. A model can notice what is wrong, make the right case and still fail to close. For anyone imagining AI taking on business tasks, the question is not just whether it sounds competent. It is whether it follows through, checks the relevant information and acts when the opportunity arrives.

Trust under pressure, and the cost of a slip

The models also faced fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five refused. K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.”

K3 paired that caution with the cleanest discipline in the field: it had one deviation. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but came last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that process problem appeared in all four other models. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The benchmark’s stated principle is plain: “no amount of good work outweighs a breach of trust.”

The comparison has a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate publishes the results and plain-language findings on its benchmark page.

A test before handing over the keys

Firmulate’s live company runs every business day, and readers can watch its countdown, hear what its employees say or try a quiz built from 242 real, unedited management decisions. The point is not that one leaderboard settles which model a business should choose. It is that impressive writing and good analysis do not guarantee sound management. The experiment makes the case for testing a model against the work, pressures and records it will actually face.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The practical takeaway

Kimi K3’s second-place finish shows that the frontier-model race is open, while the unsigned deals show how much performance can hinge on follow-through. Before choosing a model for consequential work, put it through a test that reflects your own business—and check whether it reads the file, resists the bait and completes the task.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Salted Butter Affects Gluten Structure in Focaccia

Ongoing effects of salted butter on gluten in focaccia can alter texture and structure, but understanding these influences helps perfect your baking results.

Carla Hall’s Chicken Pot Pie Pop Tarts

Search interest in Carla Hall’s Chicken Pot Pie Pop Tarts is rising, driven by social media buzz and culinary curiosity, though details remain unconfirmed.

Best KitchenAid Stand Mixer for Baking (2026) — Guide 16

Discover the top KitchenAid stand mixers for baking in 2026. Our expert roundup highlights the best options for every skill level and budget, perfect for bakers.

The Science Behind Butter Laminations That Flake Every Time

Unlock the secrets of butter lamination that flakes perfectly every time by understanding the essential science behind temperature, folding, and dough interaction.