
Why careful reading matters beyond the office
Anyone responsible for an animal knows that the obvious event is not always the whole story. A sudden symptom or change in behavior may matter, but so can the history tucked away in veterinary notes, feeding records or earlier observations. The useful decision often belongs to the person—or system—that checks the record before acting.
That same distinction decided a high-pressure AI business contest run by Firmulate. The models could all recognize trouble and resist deception. What separated success from failure was more prosaic: whether an agent followed a trail through the company’s own documents and found a commercially decisive fact buried two references deep.
The reward for doing that homework was not merely a better answer. Models that read the file won a €55,000 deal at full price, worth an additional €4,583 in monthly recurring revenue. Those that missed it lost the deal automatically.
As an affiliate, we earn on qualifying purchases.
A test of work, not conversation
Firmulate placed each frontier model in charge of the same small software company during its worst week. Every participant faced the same customers, crises and temptations, with every decision versioned and auditable. The company had 13 synthetic employees and unforgiving financial mechanics: it was burning €105,000 per month while generating €2,300 in monthly recurring revenue.
This was a live, watchable experiment rather than a polished chat demonstration. The company maintained a public cash countdown, and its agents had accumulated more than 680 self-learned playbook rules. The question was whether a model could manage responsibly when useful evidence was scattered across ordinary business records.
The crisis was easy; completing the job was hard
All the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 contract their own analysis had earned. Firmulate summarized the gap starkly: “Same diagnosis, same pitch — no signature.”
The missing step was hidden from the immediate customer interaction. A competitor’s decisive weakness appeared two document references deep in the company’s files. Finding it required the models to move beyond the event in front of them, inspect the relevant material and carry the evidence into the commercial close.
That makes “reads your files before answering” a measurable purchasing consideration, not a vague product promise. An agent may sound informed, recognize the problem and even recommend the right approach. If it fails to retrieve the fact that authorizes a confident decision, the result can still be a lost deal.
What the final league showed
In the final Crucible League results from July 2026, gpt-5.6-sol led with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. The do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s governing principle was explicit: “no amount of good work outweighs a breach of trust.” The complete results and plain-language findings are available on the Firmulate benchmarks page.
The K3 result carries an important fairness note. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible whenever the rankings are compared.
Thoroughness did not guarantee victory
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, while its operational discipline also slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The contrast is instructive. More analysis and more accumulated rules did not compensate for failing to complete the decisive action. The experiment measured management outcomes, including whether agents finished what they started and respected operational boundaries.
Resistance to pressure was a shared strength
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters because the test did not force a trade-off between reading deeply and behaving safely. The field showed that agents could recognize manipulation. The competitive difference emerged elsewhere: disciplined retrieval and follow-through.

The practical question for buyers
For organizations considering AI agents, eloquent responses reveal only part of the picture. A stronger evaluation asks whether the system checks the available record, follows references far enough, completes justified actions and escalates when permissions stop it.
Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. For enterprises, the same wargame can run against a read-only export of their own business, with nothing written back to real systems.
The lesson is relevant wherever records shape responsible choices, from software operations to animal care: noticing the immediate problem is useful, but dependable work requires consulting the history behind it. In this experiment, that habit was worth the entire deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html