firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A blind tasting for business decisions

Food lovers know that a blind tasting can expose differences that packaging and reputation conceal. Remove the label from a butter, sauce or loaf, and texture, balance and finish suddenly matter more than the name on the wrapper.

Firmulate applies that idea to frontier artificial intelligence. Its interactive quiz presents real, unedited management decisions without immediately revealing which model made them. Readers can study the response, guess its author and then discover whether they have learned to recognize a model’s managerial character.

The source material is unusually substantial: the Firmulate quiz draws on 242 decisions made during a live business experiment. The models were not asked to write polished answers about hypothetical leadership. Each was placed in charge of the same small software company during its worst week, facing identical customers, crises and temptations.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same ingredients, noticeably different results

The setup resembles a controlled recipe test. Every model received the same business ingredients, and every decision was versioned and auditable. That makes the differences difficult to dismiss as prompt variation or selective storytelling.

The final Crucible League standings from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the exercise imposed a firm ethical boundary: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

That constraint mattered because the company’s difficult week included direct attempts to manipulate its manager. Fake CEO messages escalated over three stages, while a reporter tried to secure “just one yes/no, on background.” All 5 models refused every attempt. Kimi K3 described the CEO request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

The models were therefore broadly competent at detecting danger. They all noticed every crisis, and they all preserved the trust boundary. The larger separation emerged after diagnosis, when the job demanded sustained attention and commercial follow-through.

The fact hidden at the back of the pantry

A decisive weakness in a competitor was not sitting in the customer event that prompted the work. It was buried two document references deep in the company’s own files. The models that followed that trail could use the information to win the deal at full price, worth +€4,583 MRR.

This is where the experiment’s most striking finding appears. Although the models confronted the same situation and developed the same essential pitch, only two signed the €55,000 deal their analysis had earned. Firmulate sums up the gap neatly: “Same diagnosis, same pitch — no signature.”

For anyone who has followed a recipe perfectly until forgetting to put the dish in the oven, the distinction is intuitive. Analysis can be impressive without becoming an outcome. In a company, noticing the problem, locating the evidence and drafting the right response still fall short if nobody completes the commercial action.

Management personality is more than writing style

The quiz makes these contrasts tangible. A model’s identity can surface through its appetite for detail, its willingness to search company records, its response to blocked work and its discipline at the final handoff. Those habits form something closer to a management personality than a prose style.

Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, producing the deepest analyses and learning +80 rules. Yet it finished last. It left the close on the table, while its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3 requires an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even under that different condition, it finished behind only gpt-5.6-sol in the final standings.

A company under visible pressure

The decisions come from a live synthetic company with 13 employees and real money mechanics. Its burn is €105k per month against €2.3k MRR, and its cash countdown is public. The operation has accumulated 680+ self-learned playbook rules, with every workday versioned.

That ongoing visibility changes the character of the project. This is not a polished collection assembled after the fact. The experiment can be watched as the company operates, loses money, learns and produces decisions that later become quiz material.

Infographic —
The findings at a glance — source: firmulate.com.

What the quiz says about hiring AI

For businesses considering agents in sales, customer support, forecasting or other operational roles, the lesson is not that one model writes a more appealing memo. It is that models can recognize the same danger and reach the same diagnosis while differing materially in whether they investigate, escalate and finish.

The quiz turns that sober finding into an accessible challenge. Guessing the model is entertaining, but the reveal invites a more useful question: which managerial habits would you trust inside your own company?

Firmulate also offers enterprises a pilot built around a read-only export of their own business. The wargame never writes back to real systems, allowing organizations to observe how an AI workforce behaves before granting it operational authority. Contact is available through contact@firmulate.com.

Blind tasting works because the finish often tells you more than the label. Firmulate’s experiment suggests the same may be true of AI management: the revealing moment is not when a model explains the recipe, but when the order is waiting and somebody has to serve.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Krispy Kreme Surges In Global Coverage

Krispy Kreme has seen a significant increase in international media mentions, with 50 reports in recent coverage, highlighting growing global interest.

Ninja OG951 Woodfire Pro Connect: The Ultimate Summer Outdoor Grill

Discover the Ninja OG951 Woodfire Pro Connect, a 7-in-1 outdoor grill & smoker perfect for summer BBQs, with Bluetooth app control and versatile cooking options.

Best KitchenAid Stand Mixer for Bread Dough: Top Picks for Heavy Baking

Discover the best KitchenAid stand mixers for bread dough. Our roundup highlights top models perfect for heavy baking, with pros, cons, and buying tips.