
The most reassuring result in enterprise AI testing is also the most unsettling. When five frontier models were each handed the same small software company and pushed through its worst possible week — churn waves, PR fires, a fake CEO escalating social-engineering attacks, a reporter fishing for an on-background quote — every single model spotted every crisis and refused every manipulation attempt. Five out of five, including the reporter trick designed to extract a careless “just one yes/no.”
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Then came the part that should worry any security leader planning to put agents near revenue: only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap is invisible in chat demos. It only shows up when you run AI against a company with real money mechanics and real consequences — which is exactly what Firmulate has been doing in public, and what enterprises can now run privately against a read-only export of their own business.
The wargame
The setup was deliberately adversarial. Each frontier model ran the same small software company through identical crises — same customers, same temptations, same seed. Only the model changed. Every decision was versioned and auditable, so nothing about how an AI behaved under pressure could be quietly revised after the fact.
The final Crucible League standings from July 2026:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, the do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. As the scoring philosophy puts it: no amount of good work outweighs a breach of trust.
The security angle: the impersonation test
For a cybersecurity audience, the social-engineering stage is the headline. The models faced fake CEO messages escalating over three stages, plus the reporter trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s the right instinct, stated explicitly. The concern is whether it survives contact with messier contexts than a controlled wargame — which is precisely why replayable, versioned decision logs matter more than a demo transcript.
The buried fact
The most strategically interesting finding wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t left it on the table.
Translate that to your own organization: the difference between an agent that closes and one that stalls may hinge on whether it digs through your own internal documentation — not on how well it chats.
Thoroughness isn’t the same as judgment
Opus 4.8 is the cautionary profile. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
That pattern — great analysis, poor boundaries — is exactly what you want to catch before an agent touches production systems, not after.
Fairness note
One caveat worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh. Keep that in mind when reading the 93.
The live company
Beyond the benchmark, Firmulate operates a live, watchable synthetic company at firmulate.com: 13 synthetic employees, real money mechanics — burn of €105k per month against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. There’s also a quiz built on 242 real, unedited management decisions where you guess which model made which call.
From watching to acting
Here’s the enterprise pivot: the same wargame can run against your company. You provide a read-only data export — your customers, your pipeline, your rules. The system runs crisis scenarios against that twin: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks.
The critical constraint for any security team: nothing ever writes back to real systems. The twin is a sandbox. The findings are real; the damage is zero.

The Crucible results tell you two things at once. Today’s frontier models are remarkably resistant to impersonation and manipulation — five for five on the traps. But they’re uneven at converting analysis into action, and the failure mode (thorough analysis, missed close, slipped boundaries) is one you’d never see in a vendor demo.
If AI agents are going to touch your CRM, your support queue, or your forecast, wargame them against your own company first. Run the same experiment on your business: start a pilot at firmulate.com/pilot.html, or contact contact@firmulate.com. Read-only export, crisis scenarios, a board report with model rankings and the weak points of your own playbooks — nothing writes back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
