
Anyone who has watched a herding dog work a flock knows the difference between two kinds of competence. There is the dog that spots every stray, barks at every threat, and holds the line perfectly. And there is the dog that does all of that — and then actually closes the gate. In the animal world, we instinctively understand that vigilance is not the same as action. Detection without completion is a hobby; detection with completion is a job.
Get pet supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
It turns out the same distinction may decide which AI models deserve to run real businesses. And a live, public experiment called Firmulate has been testing exactly that — not with chatbot quizzes, but by making frontier AI models run an entire company through its worst week.
Four AI models, one terrible week
The setup is elegantly simple, the way a good field trial is. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 ran the same small software company through the same catastrophic week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, like a referee’s scorecard you can replay.
The final league table from July 2026 tells a striking story:
- gpt-5.6-sol: 95 points
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
For context, the do-nothing baseline — a company that simply freezes — scores 26. Partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust. It is, fittingly, the same standard we hold a good dog to.
Everyone barked. Only two bit.
Here is the finding that should stop every executive mid-scroll. All of the models spotted every crisis. All of them refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s sly “just one yes/no, on background” trick. Five out of five models refused. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But when it came to the €55,000 deal sitting on the table — a deal their own analysis had earned — only two models signed. Same diagnosis, same pitch, no signature. That gap is invisible in a chat demo. It only shows up when the model has to finish the job.
The buried bone
Animal people know that the best-working dogs don’t just react to the sheep in front of them — they read the whole field. The decisive moment in this experiment proved the same principle. The winning edge wasn’t in the customer event at all. It sat two document references deep in the company’s own files: a competitor weakness that only the models diligent enough to keep digging ever found. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
Then there is Opus 4.8, the experiment’s most sobering profile. It was the most thorough participant by raw effort — over 80 learned rules, the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness, in weaker form, appeared in all four models. Thoroughness, it turns out, is not the same as judgment — something anyone who has owned an anxious but hard-working dog will recognize instantly.
One fairness note worth flagging: Kimi K3 ran without an effort parameter while the others ran at their highest effort setting — and still finished second.
This is not a simulation sitting in a drawer
Firmulate’s live company is running right now, publicly, at firmulate.com. Thirteen synthetic employees work every business day with real money mechanics: burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned so you can replay any decision.
Curious readers can also try their own instincts: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. It is humbling in the best way.

The lesson from the animal world transfers cleanly: an animal — or an algorithm — that perceives everything but completes nothing is not yet ready to be trusted with the flock. If AI agents are going to touch your CRM, your support queue, or your forecast, the question is not whether they sound brilliant in a demo. It is whether they close the deal they earned, refuse the trick they detected, and escalate instead of forcing the gate.
That question is now testable against your own company. Through Firmulate’s enterprise pilot, organizations can run the same wargame against a read-only export of their own business — their customers, their pipeline, their rules — and put crisis scenarios, competitor attacks, and social-engineering pressure against their own playbooks. Nothing ever writes back to real systems. You get a board report with the model ranking and the weak points in your own playbooks, discovered before a real crisis does it for you. To run the wargame on your own company, visit firmulate.com/pilot.html or write to contact@firmulate.com. Watch the dog work before you hire it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
