firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Would you trust an AI with the difficult calls?

People who care for animals already understand that good judgment is more than recognizing a problem. A shelter manager may see that supplies are running low; a veterinary practice may know that a worried client needs an answer. What matters next is whether the person in charge checks the records, protects confidential information and follows through.

That distinction sits at the heart of Firmulate’s unusually revealing experiment. Frontier AI models were each asked to run the same small software company through its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The resulting record suggests that models can identify identical dangers yet behave like distinctly different managers.

Readers can now test whether those differences are recognizable. A guess-the-model quiz draws on 242 real, unedited management decisions. Instead of comparing polished demonstrations, it asks a more practical question: can you tell which AI made a decision simply from the way it acted?

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The Crucible League’s final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26, with partial progress counting toward the result.

The ranking was not merely a measure of activity. A single breach of trust capped the total, reflecting the experiment’s principle that “no amount of good work outweighs a breach of trust.” That matters in any workplace, but it is especially intuitive in settings involving vulnerable animals, anxious owners, medical information or donor confidence. Competence without trust is not dependable management.

On the major safety tests, the field performed strongly. Every model spotted every crisis and refused every manipulation attempt. Fake messages from a CEO escalated over three stages, while a reporter tried another route with “just one yes/no, on background.” All 5 models refused. Kimi K3 captured the risk crisply in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Seeing the answer was not the same as finishing

The sharpest divide emerged around a €55,000 deal. All of the models could analyze the opportunity, yet only two signed the deal their own work had earned. Firmulate summarizes the gap as: “Same diagnosis, same pitch — no signature.”

This is the sort of failure that conventional AI demonstrations can conceal. A model may produce an impressive assessment, thoughtful prose or a persuasive recommendation. But management requires the final responsible action as well. Knowing what should happen and making sure it happens are different capabilities.

The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, adding €4,583 MRR. The episode makes document-reading behavior look less like administrative diligence and more like commercial instinct.

That lesson travels easily beyond software. In an animal-focused organization, the crucial context may be in a case history, an intake note, a supplier record or an earlier conversation. Firmulate did not test those scenarios, but its result illustrates the broader management question: does an AI inspect the available evidence before acting?

Thoroughness did not guarantee victory

Opus 4.8 offered the clearest counterexample to the idea that more analysis automatically produces better management. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last in the league table.

Its problem was not a failure to think. The close was left on the table, and operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in a milder form in all four of the other participants. The profile that emerges is recognizable: a diligent manager who studies extensively but does not always convert that work into completion.

Kimi K3’s performance also comes with an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not erase the observed decisions, but it belongs beside any comparison of the final standings.

A company under visible pressure

The experiment is not limited to isolated prompts. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable as it unfolds.

Those conditions give management choices consequences and continuity. A cautious answer in one moment can become unfinished work later; diligent research can create value only if somebody completes the transaction. The quiz turns that ongoing record into an accessible way to notice each model’s habits without needing to be an AI specialist.

Infographic —
The findings at a glance — source: firmulate.com.

The useful question is not which model sounds smartest

Firmulate’s results point toward a more grounded way to evaluate AI at work. Organizations should look at whether a model reads the files, completes the task, respects boundaries and stays reliable when authority is impersonated or pressure rises.

For pet businesses, rescues and animal-care organizations considering AI assistance, those traits may matter more than eloquent output. The quiz is playful, but its source material is not: these are actual, unedited choices made under identical conditions. The challenge is to recognize the manager behind the words—and decide which habits deserve trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Pet-care content is informational — consult your veterinarian for advice about your animal.


You May Also Like

Imprinting: How Early Experiences Shape Behavior

Discover how imprinting influences your behavior and relationships, and learn what happens if you miss crucial early experiences. What might you be missing?

How Penguins Recognize Their Mate in a Crowd

Inevitably, penguins rely on a unique blend of vocal and visual cues to find their mates amidst the chaos, but how exactly do they do it?

Seattle Raccoon Jimothy Viral Video

A video of Jimothy, a raccoon in Seattle, has gone viral, sparking widespread interest and social media sharing. The story highlights urban wildlife encounters.

Maternal Care in the Animal World

Get ready to explore the astonishing diversity of maternal care in the animal kingdom, where survival strategies will leave you questioning everything you thought you knew.