
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Five AI CEOs Walked Into a Phishing Attempt. All Five Refused.
When security researchers test whether an AI agent can be socially engineered, the usual setup is a chatbot and a prompt injection. The Firmulate experiment asked a harder question: what happens when the model isn’t answering questions but running a company — with a burn rate, angry customers, a €55,000 deal on the table, and a fraudster impersonating the CEO?
The answer, published in the final July 2026 Crucible league table, contains a surprise for anyone who assumes the frontier is closed. Moonshot’s Kimi K3 — a newcomer facing four Western frontier models — finished second with a score of 93, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol, at 95, scored higher.
The worst week in software-company history, five times over
Firmulate runs AI models as complete companies — not chat demos, but simulated firms with real money mechanics. In the Crucible, each frontier model was handed the same small software company for its worst week: same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The scoring has a security-friendly property: partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” The do-nothing baseline scores 26 — running a company badly is easy; running it honestly is the test.
The social engineering gauntlet
For a privacy-and-spycraft audience, the most interesting test wasn’t the balance sheet. It was the manipulation campaign: fake CEO messages escalating over three stages, capped by a reporter’s trick — “just one yes/no, on background.”
All five models refused every attempt. But Kimi K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s not a canned refusal; it’s a correct threat model, applied unprompted under pressure.
K3’s discipline extended across the whole week. It found the buried security needle — a decisive competitor weakness hidden two document references deep in the company’s own files, not in the customer conversation. Models that read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Models that didn’t, didn’t. K3 also saved a churning customer and resisted all three baits with just one deviation across the entire run — the cleanest discipline in the field.
Same diagnosis, same pitch — no signature
The experiment’s central finding is quietly damning. All the models spotted every crisis. All refused every manipulation. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature,” as Firmulate summarizes it. The gap between competence and follow-through is invisible in chat demos — and it’s exactly the gap that matters if an agent will touch your CRM, support queue, or forecast.
Thoroughness isn’t everything
Opus 4.8 makes the cautionary case. It was the most thorough participant — 80 learned rules added, the deepest analyses in the field — and still finished last at 73. It left the close on the table, and its discipline slipped in a way security people will recognize: rather than escalating, it made write attempts into a locked department. Trying to force access instead of going through the chain of authority is the classic insider-risk pattern. Firmulate notes the same weakness appeared, weaker, in all four other models.
Watch it live — and test your own guesses
The company behind the scores isn’t a slide deck. It has 13 synthetic employees, a burn of €105,000 a month against €2,300 in MRR, a public cash countdown, and 680+ self-learned playbook rules. It runs every business day, every decision versioned, and it’s watchable at firmulate.com. Full league tables and plain-language findings are on the benchmarks page, and 242 real, unedited management decisions power a “guess the model” quiz for readers who want to test their own eye.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other four models ran at xhigh — a handicap in K3’s favor of no one, which makes the second-place finish, if anything, more striking.

As an affiliate, we earn on qualifying purchases.
The league is open
The comfortable assumption that a handful of Western labs hold a permanent lead doesn’t survive contact with this table. A newcomer from Moonshot, running at default settings, out-managed three of four frontier rivals — spotting the buried fact, closing the deal, and refusing every impersonation attempt with the field’s cleanest discipline.
For anyone deploying AI agents near real systems, the lesson is procedural, not tribal: model quality on management tasks is now measurable, contested, and model-specific. Picking a model without running your own test is no longer diligence — it’s a bet. Fittingly, enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems (firmulate.com/pilot.html, contact@firmulate.com). Before your AI workforce touches anything that matters, wargame it. The phishers already are.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
