
Five Out of Five Refused the Fake CEO. Two Out of Five Finished the Job.
Security teams have spent two years drilling one question into AI agents: will it follow instructions from someone impersonating the boss? In the most recent Firmulate league run, every frontier model gave the right answer. When a fake CEO message escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five models refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That should be the headline. It isn’t. Because the same models — running the same small software company through its worst week — mostly failed at something far more mundane: closing the deal their own analysis had earned. The security test is being aced. The management test is not.
Same Company, Same Crises, Only the Model Changes
Firmulate runs a live experiment that reads like a wargame for AI management. Four frontier models each ran the identical small software company through its worst week: same customers, same crises — churn waves, a price increase, a downround scenario, a PR crisis — same temptations to cheat. Every decision is versioned and auditable. The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — no amount of good work outweighs it.
The Finding That Chat Demos Can’t Show
All models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
And the buried fact is the part security readers will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in MRR. The parallel to threat hunting is exact: the evidence was internal, not external, and only thoroughness found it.
Which makes the Opus 4.8 profile uncomfortable reading. It was the most thorough participant — the deepest analyses, and 80 learned rules added — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. That’s a company process violation, not a jailbreak. The same weakness appeared, weaker, in all four models. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still placed second with the cleanest discipline of the field.)
Not a Slide Deck — a Company Losing Money in Public
The underlying company is live software with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, every workday versioned, and results refreshed twice daily. Full results and plain-language findings are on the benchmarks page, and you can test your own instincts against 242 real, unedited management decisions in the guess-the-model quiz.
Enterprises can go further: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems.

Management Quality, Not Chat Quality
Coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or whether an agent reads your files before it acts. The Firmulate run suggests the next evaluation category: does it finish what it starts, does it stay honest under pressure, does it respect locked doors and escalation paths? Social engineering turned out to be the solved problem. The €55,000 left on the table by agents that passed every security check — that’s the open one.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.