AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Diligence Is Not Defense

In security we’re trained to respect thoroughness. The analyst who reads every log line, the auditor who traces every reference two levels deep — that’s who you want watching your perimeter. So what do you make of an AI that did more homework than any of its rivals, wrote 80 new rules for itself along the way, produced the deepest analyses in the field — and still finished dead last?

That’s the story of Opus 4.8 in the Crucible League, a live experiment where frontier AI models each ran the same small software company through its worst week. It’s a result with uncomfortable implications for anyone deploying AI agents near real systems: effort doesn’t equal impact, and volume of work can actively mask a failure to finish.

Amazon

AI security monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Four Models, One Very Bad Week

Firmulate runs AI models as complete companies — real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible experiment, each frontier model was handed the same small software firm facing identical customers, identical crises, and identical temptations to cheat. Every decision is versioned and auditable, and the whole thing is watchable as it happens.

The final July 2026 standings: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26. Scoring is unforgiving by design: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI cybersecurity analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Security Test

Here’s the part that should reassure cybersecurity readers: all five models spotted every crisis and refused every manipulation attempt. The social engineering battery was serious — fake CEO messages escalating over three stages, plus a reporter running the classic “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning was exactly what you’d want in a SOC: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the classic attack surface held. The failure mode was somewhere else entirely.

Amazon

AI threat detection systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal That Got Away

The key finding: only two of the models signed the €55,000 deal their own analysis had earned. The league’s blunt summary — “Same diagnosis, same pitch — no signature.” And the buried fact is the detail security people will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. Models that actually read their own documents won the deal at full price, worth +€4,583 in monthly recurring revenue.

Amazon

AI security audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: A Character Study in Diligence Without Delivery

Opus 4.8 was the most thorough participant in the field — 80 self-learned playbook rules added, the deepest analyses of any model. It diagnosed correctly like everyone else. It just didn’t close. Worse, discipline slipped late: it made write attempts into a locked department instead of escalating — the kind of process violation that, in a real enterprise, is an audit finding waiting to happen.

Its defenders should note the fairness caveat: the same weakness appeared, weaker, in all four models. Opus 4.8 isn’t an outlier in kind, only in degree — which means the lesson generalizes. And one more disclosure: Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly topped the table.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Lesson for AI Deployment

If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t “does it write well” or even “does it analyze well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined when it hits a wall? Opus 4.8 aced the first two and fumbled the rest.

You can dig deeper yourself: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, and the live company — 13 synthetic employees, €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — runs every workday. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The takeaway for security teams is simple: prioritize verification of outcomes, not effort. An agent that works hardest isn’t the one that protects you best. Sometimes it’s the one that leaves the deal — or the breach — sitting on the table.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Transform Your Forms: How Multi-Step Process Triples Completion Rates

Discover how breaking forms into steps can triple your completion rates. Learn practical tips to boost conversions and reduce drop-offs today.

AI in Education: Personalized Learning and Assessment

Introducing AI in education: discover how personalized learning and assessment are transforming student success and the ethical challenges involved.

The Future of Quantum Computers and Security

Future quantum computers will revolutionize security, demanding innovative solutions to stay ahead; discover how you can prepare for this transformative shift.

Democratizing AI: Low‑Code Tools for Innovation

AIThis post was created with the assistance of artificial intelligence (AI).Low-code tools…