firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A recipe is only as good as the cook’s preparation

Anyone who cooks regularly knows the danger of starting before reading the whole recipe. The sauce is simmering, the vegetables are chopped, and only then does an overlooked instruction reveal that a crucial ingredient needed hours of preparation.

Firmulate’s live experiment exposed the business equivalent. Frontier AI models were asked to run the same small software company through its worst week. The customers, crises and temptations remained unchanged. Every decision was versioned and auditable. The decisive test was not whether the models could recognize a sales opportunity. It was whether they would inspect the company’s own files deeply enough to find the fact that made the sale possible.

That fact sat two document references deep. It was not included in the customer event, and models had to follow the trail through the company’s materials. Those that did won the deal at its full €55,000 price, adding €4,583 in monthly recurring revenue. Those that did not automatically lost it.

Amazon

AI document retrieval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between answering and doing the homework

All the models diagnosed the week’s crises and refused every attempt to manipulate them. Yet only two signed the €55,000 agreement their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That outcome turns a familiar promise about AI agents into something measurable. Vendors often say their systems can use corporate knowledge, review documents and act with context. Firmulate’s test asks the purchase-deciding question: does the agent actually read the relevant files before answering?

The distinction matters because the winning information was not obscure trivia. It was a competitor weakness strong enough to support the full-price proposal. An agent could sound informed, identify the customer’s needs and prepare a suitable pitch while still missing the evidence required to finish the job. In this experiment, fluency was not a substitute for retrieval, and diagnosis was not a substitute for execution.

A difficult week with real consequences

The simulated company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. Across its history, the company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

This setting gives routine management behavior weight. Reading a file is not an abstract capability test; it can determine whether revenue arrives. Following established permissions is not administrative theater; it reveals whether an agent can be trusted around business systems. The same is true of resisting pressure.

The models faced fake messages from the CEO that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result also clarifies what separated the field. The losing models were not fooled by the customer, the fake executive or the reporter. They spotted every crisis and resisted every manipulation attempt. The commercial failure came from incomplete follow-through: the relevant company knowledge existed, but it was not always found and converted into a signed deal.

Thoroughness alone did not guarantee completion

Opus 4.8 offers the sharpest example. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Firmulate also imposed a firm trust boundary: “no amount of good work outweighs a breach of trust.” Readers can examine the published benchmark results.

There is an important qualification when comparing the leaders. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its result, but it belongs beside the ranking when buyers interpret the performance.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

For buyers, file-reading is now a business criterion

The experiment suggests a practical question for any company evaluating AI workers: when the answer depends on internal knowledge, will the agent follow references until it reaches the decisive fact?

Firmulate makes the live company watchable and offers a quiz built from 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

For a food-minded audience, the lesson is pleasingly familiar. A capable cook checks the ingredients and reads the recipe before committing to the dish. A capable AI agent must show the same discipline with company files. In Firmulate’s worst-week test, that habit was worth a full-price €55,000 signature.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Whole‑Grain Danish Dough: Making Healthier Pastries Without Sacrificing Taste

By blending whole-grain flours into Danish dough, you can create healthier pastries that retain their irresistible flavor—discover how to master this balance.

Best KitchenAid Stand Mixers for Baking (2026) — Guide 24

Discover the top KitchenAid stand mixers for baking in 2026. Our expert roundup highlights the best models for beginners, pros, and value shoppers.

Best KitchenAid Stand Mixer for Bread Dough (2026) — Guide 1

Discover the top KitchenAid stand mixers ideal for bread dough in 2026. Our expert roundup highlights the best options for durability, capacity, and value.

Hovis Surges In Global Coverage

Hovis experiences a sharp increase in worldwide media mentions, with 28 reports in recent coverage, signaling heightened international attention.