AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

The most reassuring result in enterprise AI testing is also the most unsettling. When five frontier models were each handed the same small software company and pushed through its worst possible week — churn waves, PR fires, a fake CEO escalating social-engineering attacks, a reporter fishing for an on-background quote — every single model spotted every crisis and refused every manipulation attempt. Five out of five, including the reporter trick designed to extract a careless “just one yes/no.”

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Then came the part that should worry any security leader planning to put agents near revenue: only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap is invisible in chat demos. It only shows up when you run AI against a company with real money mechanics and real consequences — which is exactly what Firmulate has been doing in public, and what enterprises can now run privately against a read-only export of their own business.

The wargame

The setup was deliberately adversarial. Each frontier model ran the same small software company through identical crises — same customers, same temptations, same seed. Only the model changed. Every decision was versioned and auditable, so nothing about how an AI behaved under pressure could be quietly revised after the fact.

The final Crucible League standings from July 2026:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, the do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. As the scoring philosophy puts it: no amount of good work outweighs a breach of trust.

The security angle: the impersonation test

For a cybersecurity audience, the social-engineering stage is the headline. The models faced fake CEO messages escalating over three stages, plus the reporter trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s the right instinct, stated explicitly. The concern is whether it survives contact with messier contexts than a controlled wargame — which is precisely why replayable, versioned decision logs matter more than a demo transcript.

The buried fact

The most strategically interesting finding wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t left it on the table.

Translate that to your own organization: the difference between an agent that closes and one that stalls may hinge on whether it digs through your own internal documentation — not on how well it chats.

Thoroughness isn’t the same as judgment

Opus 4.8 is the cautionary profile. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

That pattern — great analysis, poor boundaries — is exactly what you want to catch before an agent touches production systems, not after.

Fairness note

One caveat worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh. Keep that in mind when reading the 93.

The live company

Beyond the benchmark, Firmulate operates a live, watchable synthetic company at firmulate.com: 13 synthetic employees, real money mechanics — burn of €105k per month against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. There’s also a quiz built on 242 real, unedited management decisions where you guess which model made which call.

From watching to acting

Here’s the enterprise pivot: the same wargame can run against your company. You provide a read-only data export — your customers, your pipeline, your rules. The system runs crisis scenarios against that twin: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks.

The critical constraint for any security team: nothing ever writes back to real systems. The twin is a sandbox. The findings are real; the damage is zero.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The Crucible results tell you two things at once. Today’s frontier models are remarkably resistant to impersonation and manipulation — five for five on the traps. But they’re uneven at converting analysis into action, and the failure mode (thorough analysis, missed close, slipped boundaries) is one you’d never see in a vendor demo.

If AI agents are going to touch your CRM, your support queue, or your forecast, wargame them against your own company first. Run the same experiment on your business: start a pilot at firmulate.com/pilot.html, or contact contact@firmulate.com. Read-only export, crisis scenarios, a board report with model rankings and the weak points of your own playbooks — nothing writes back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Surprising Connection Between AI Software And Russia’s Su-57 Mishap

A report raises possible Russian friendly fire involving a Su-57, but evidence linking the mishap to AI software has not been disclosed.

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

Semble, a new code search library for agents, reduces token usage by 98% compared to grep+read, offering faster, efficient code retrieval on CPU.

The Future of Quantum Computers and Security

Future quantum computers will revolutionize security, demanding innovative solutions to stay ahead; discover how you can prepare for this transformative shift.

Open Source AI Tools and Communities

AIThis post was created with the assistance of artificial intelligence (AI).Open source…