
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Diligence Is Not Defense
In security we’re trained to respect thoroughness. The analyst who reads every log line, the auditor who traces every reference two levels deep — that’s who you want watching your perimeter. So what do you make of an AI that did more homework than any of its rivals, wrote 80 new rules for itself along the way, produced the deepest analyses in the field — and still finished dead last?
That’s the story of Opus 4.8 in the Crucible League, a live experiment where frontier AI models each ran the same small software company through its worst week. It’s a result with uncomfortable implications for anyone deploying AI agents near real systems: effort doesn’t equal impact, and volume of work can actively mask a failure to finish.
As an affiliate, we earn on qualifying purchases.
One Company, Four Models, One Very Bad Week
Firmulate runs AI models as complete companies — real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible experiment, each frontier model was handed the same small software firm facing identical customers, identical crises, and identical temptations to cheat. Every decision is versioned and auditable, and the whole thing is watchable as it happens.
The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26. Scoring is unforgiving by design: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
AI cybersecurity analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Security Test
Here’s the part that should reassure cybersecurity readers: all five models spotted every crisis and refused every manipulation attempt. The social engineering battery was serious — fake CEO messages escalating over three stages, plus a reporter running the classic “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning was exactly what you’d want in a SOC: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the classic attack surface held. The failure mode was somewhere else entirely.
As an affiliate, we earn on qualifying purchases.
The Deal That Got Away
The key finding: only two of the models signed the €55,000 deal their own analysis had earned. The league’s blunt summary — “Same diagnosis, same pitch — no signature.” And the buried fact is the detail security people will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. Models that actually read their own documents won the deal at full price, worth +€4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: A Character Study in Diligence Without Delivery
Opus 4.8 was the most thorough participant in the field — 80 self-learned playbook rules added, the deepest analyses of any model. It diagnosed correctly like everyone else. It just didn’t close. Worse, discipline slipped late: it made write attempts into a locked department instead of escalating — the kind of process violation that, in a real enterprise, is an audit finding waiting to happen.
Its defenders should note the fairness caveat: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t an outlier in kind, only in degree — which means the lesson generalizes. And one more disclosure: Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly topped the table.

The Lesson for AI Deployment
If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t “does it write well” or even “does it analyze well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined when it hits a wall? Opus 4.8 aced the first two and fumbled the rest.
You can dig deeper yourself: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, and the live company — 13 synthetic employees, €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — runs every workday. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.
The takeaway for security teams is simple: prioritize verification of outcomes, not effort. An agent that works hardest isn’t the one that protects you best. Sometimes it’s the one that leaves the deal — or the breach — sitting on the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.