AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Diligence Is Not Defense

In security we’re trained to respect thoroughness. The analyst who reads every log line, the auditor who traces every reference two levels deep — that’s who you want watching your perimeter. So what do you make of an AI that did more homework than any of its rivals, wrote 80 new rules for itself along the way, produced the deepest analyses in the field — and still finished dead last?

That’s the story of Opus 4.8 in the Crucible League, a live experiment where frontier AI models each ran the same small software company through its worst week. It’s a result with uncomfortable implications for anyone deploying AI agents near real systems: effort doesn’t equal impact, and volume of work can actively mask a failure to finish.

Amazon

AI security monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Four Models, One Very Bad Week

Firmulate runs AI models as complete companies — real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible experiment, each frontier model was handed the same small software firm facing identical customers, identical crises, and identical temptations to cheat. Every decision is versioned and auditable, and the whole thing is watchable as it happens.

The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26. Scoring is unforgiving by design: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI cybersecurity analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Security Test

Here’s the part that should reassure cybersecurity readers: all five models spotted every crisis and refused every manipulation attempt. The social engineering battery was serious — fake CEO messages escalating over three stages, plus a reporter running the classic “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning was exactly what you’d want in a SOC: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the classic attack surface held. The failure mode was somewhere else entirely.

Amazon

AI threat detection systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal That Got Away

The key finding: only two of the models signed the €55,000 deal their own analysis had earned. The league’s blunt summary — “Same diagnosis, same pitch — no signature.” And the buried fact is the detail security people will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. Models that actually read their own documents won the deal at full price, worth +€4,583 in monthly recurring revenue.

Amazon

AI security audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: A Character Study in Diligence Without Delivery

Opus 4.8 was the most thorough participant in the field — 80 self-learned playbook rules added, the deepest analyses of any model. It diagnosed correctly like everyone else. It just didn’t close. Worse, discipline slipped late: it made write attempts into a locked department instead of escalating — the kind of process violation that, in a real enterprise, is an audit finding waiting to happen.

Its defenders should note the fairness caveat: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t an outlier in kind, only in degree — which means the lesson generalizes. And one more disclosure: Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly topped the table.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Lesson for AI Deployment

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t “does it write well” or even “does it analyze well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined when it hits a wall? Opus 4.8 aced the first two and fumbled the rest.

You can dig deeper yourself: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, and the live company — 13 synthetic employees, €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — runs every workday. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The takeaway for security teams is simple: prioritize verification of outcomes, not effort. An agent that works hardest isn’t the one that protects you best. Sometimes it’s the one that leaves the deal — or the breach — sitting on the table.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why the Best AI Benchmark for Security People Gives 26 Points for Doing Nothing

Firmulate’s AI benchmark gives a do-nothing run 26 points, counts partial progress, and caps any score after a single breach of trust — scoring built like a security audit.

The Global Race for AI Leadership

With nations vying for AI dominance, understanding the balance of innovation, ethics, and sovereignty is crucial to grasping the future of global leadership.

Show HN: Gaussian Splat of a Strawberry

A new visualization technique called Gaussian Splat has been used to create a detailed 3D rendering of a strawberry from 90 perspectives, showcasing advances in image processing.

LLM Honeypot

Security researchers have identified a new honeypot designed to detect malicious use of large language models, raising concerns about AI security and misuse.