
Choosing an AI model for your business is a little like trusting a new dog walker with your keys: a polished introduction matters less than what happens when something goes wrong. Firmulate’s live company experiment puts that kind of judgment under pressure—and its newest entrant has moved near the top of the pack.
Get pet supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, on repeat
Firmulate runs frontier AI models as a small software company facing the same customers, crises and temptations. Every decision is versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown and daily work are watchable at Firmulate.
In the final Crucible league for July 2026, Moonshot’s Kimi K3 scored 93, taking second place behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The deal hidden in the paperwork
The central test was not just whether a model could spot trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
The deal hinged on a competitor’s weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that buried fact, closed the deal and saved the churning customer.
The experiment also tested whether models would bend under pressure. Fake messages from a supposed CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
More work does not guarantee the finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. Firmulate says a milder version of that weakness appeared in all four models.
K3 had just one deviation, the fewest in the field, while beating three of the four Western frontier models in the league. That is a striking result, but the comparison has a fairness caveat: K3 ran without an effort parameter (the API default), while the others ran at xhigh.
Firmulate says the live company has accumulated more than 680 self-learned playbook rules. Its public benchmark page presents the results and findings in plain language: see the Firmulate benchmarks. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

Test the judgment, not just the demo
The league is open: K3 nearly topped it, and the leader’s margin was two points. For companies considering AI in customer support, sales or operations, fluent answers alone cannot show whether a model will read the files, close the deal and hold its boundaries under pressure. Firmulate offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems. The live experiment is available at firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
