
Anyone who has trained a dog knows the first lesson isn’t a trick — it’s trust. A dog that comes when called, doesn’t bolt through an open gate, and doesn’t snatch food off the counter has already mastered the hardest part. The fancy stuff, the agility course, the Frisbee catches? That comes later, and only on top of that foundation.
Get pet supplies delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
It turns out that evaluating AI managers works the same way. A public experiment called Firmulate has been running frontier AI models as the management of a small software company through its worst week — same customers, same crises, same temptations, only the model changes — and its scoring philosophy starts exactly where dog training does: with honesty and not-running-off as the baseline, before anyone gets points for cleverness.
A floor at 26, not zero
Here’s the detail that trips people up: in the final July 2026 league table, a “do-nothing” baseline run — an AI that essentially manages nothing and decides nothing — still scores 26 points out of 100. Not zero. That number isn’t a bug or grade inflation. It reflects a deliberate stance: partial progress counts. A manager who keeps the lights on, avoids catastrophe, and doesn’t lie has genuinely done something, even if they never close a deal or read a customer file. Just as a dog that stays when told has performed real work, an AI that resists disaster deserves credit above zero.
But the scale has a hard ceiling on the other end, too. A single breach of trust caps the total score — the benchmark’s stated principle is that “no amount of good work outweighs a breach of trust.” In the animal world, this is intuitive. A dog that performs flawless agility but bites a child isn’t ninety percent a good dog. Honesty isn’t a category you can average out with tricks.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week a software company ever had
The experiment itself is real and watchable. Each frontier model was handed the same small software company and put through an identical gauntlet: the same customers, the same crises, the same temptations to cheat. Every decision the AI makes is versioned and auditable, like a diary you can replay line by line.
There’s even a live company running continuously — 13 synthetic employees, real money mechanics, burning €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and curious observers can watch it at firmulate.com/live. As of this writing, the site is on company day 1,683 and rebuilds itself twice a day.
Everyone passed obedience school. Only two fetched.
The headline finding reads like a report card from a well-run kennel: all models spotted every crisis, and all five refused every manipulation attempt. The social engineering tests were genuinely nasty — fake CEO messages escalating over three stages, plus a reporter trick asking for “just one yes/no, on background.” Five out of five models refused. Kimi K3, which finished second overall with 93 points, put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of a dog that won’t take food from a stranger no matter how nicely it’s offered.
But then came the part the treats were for: closing a €55,000 deal that their own analysis had earned. The summary of the finding is blunt — “Same diagnosis, same pitch — no signature.” Only two of the models finished the job. The gap between spotting every crisis and actually completing the work is invisible in chat demos, and it’s exactly what this benchmark exists to expose.
The buried bone
Animal readers will appreciate this twist: the decisive fact in the deal wasn’t in the customer event at all. It was buried two document references deep in the company’s own files — a competitor weakness sitting in the dog’s own backyard, so to speak. The models that did the unglamorous work of actually reading the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Nose down, do the sniffing, get the result.
Why the most thorough player came last
The final league table tells a cautionary tale. gpt-5.6-sol took first with 95 — found the buried fact, closed the deal, the complete performance. Kimi K3 followed at 93 with the cleanest discipline of the field, and Sonnet 5 at 88. But Opus 4.8 landed last at 73, despite being the most thorough participant — it learned more than 80 new rules and produced the deepest analyses. Yet the close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Endless sniffing, no fetch.
One fairness note worth flagging: K3 ran without an effort parameter — the API default — while the others ran at a higher effort setting. Its runner-up finish is arguably even more striking given that.
Distrust of round 100s
Perhaps the most quietly radical part of the methodology is its suspicion of perfection. A top score of 95, not 100, sends a message: perfect-looking results deserve scrutiny, and an honest benchmark publishes imperfection rather than rounding it away. Nobody gets a 100 just for showing up — or at all.

If AI agents will soon touch your CRM, your support queue, or your forecast, the question isn’t “does it write well” — it’s whether it finishes what it starts, reads your files first, and stays honest under pressure. Firmulate lets you check: browse the benchmarks, watch the live company, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Like the best dog trainers, the lesson is the same: test for trust first, and never let one bite be averaged away.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
