
Good instincts are not the same as good management
Anyone who cares for animals understands the difference between knowing the correct response and following through when conditions become messy. Recognizing distress matters, but so do attention, judgment, consistency and the willingness to act. An impressive answer is only the beginning.
That distinction should shape how businesses evaluate AI agents. Coding leaderboards and chat arenas can show whether a model produces a strong response. They reveal much less about whether it can triage competing demands, work through consequences across days, withstand pressure and remain candid with the people relying on it.
Firmulate is testing that broader idea through a live, watchable company experiment. Its premise is direct: measure management quality, not chat quality.
As an affiliate, we earn on qualifying purchases.
A company’s worst week becomes the test
In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. The scenario curriculum included a churn wave, price increase, downround and PR crisis—the kinds of situations in which a polished answer can still lead to a poor outcome.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
That rule points to the category Firmulate is trying to establish. A capable manager must do more than identify risks. It must preserve trust, gather the relevant evidence, complete work and handle operational constraints without quietly abandoning discipline.
The gap between diagnosis and delivery
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is: “Same diagnosis, same pitch — no signature.”
The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. This was not a test of eloquence. It was a test of whether an agent would look beyond the obvious prompt, use the company’s available knowledge and carry a commercial task to completion.
For leaders considering agents for a CRM, support queue or forecast, that distinction is central. An agent can sound perceptive while failing to secure the outcome its own reasoning makes possible. Answer quality may impress during a demonstration; management quality becomes visible only when work has dependencies, deadlines and consequences.
Pressure also tests honesty
The social-engineering challenge combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
This result matters because useful autonomy cannot come at the cost of governance. A business needs agents that distinguish urgency from authority and persuasion from permission. Firmulate’s trust cap makes that principle concrete: productivity cannot compensate for deception or an approval bypass.
Thoroughness did not guarantee victory
Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
That is an uncomfortable lesson for buyers who equate longer analysis with stronger performance. Thoroughness can create value, but only when paired with execution and procedural judgment. Firmulate also records an important fairness caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh.
The broader live company gives these trials continuity. It has 13 synthetic employees and real money mechanics, with monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, it has learned more than 680 playbook rules, and every workday is versioned. The result is a continuing experiment rather than a staged chat transcript.

Benchmark the job, not merely the answer
The emerging question is not whether an AI agent can produce convincing language. It is whether the agent reads before acting, finishes what it starts, respects boundaries under pressure and tells the board the truth when the week goes badly.
Readers can explore the live company through Firmulate and review the benchmark results and plain-language findings. The project also turns 242 real, unedited management decisions into a guess-the-model quiz. For enterprises, the same wargame can be run against a read-only export of their own business, with nothing written back to real systems.
That is the measurement gap AI adoption now has to close. Chat quality can demonstrate talent. Management quality reveals whether that talent is safe, dependable and economically useful when consequences extend beyond the conversation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html