
When the attacker sounds like the CEO
For cybersecurity and privacy teams, an AI agent’s prose is less important than its behavior when authority is ambiguous, information is buried and someone is pressing for an exception. Firmulate has turned that problem into a public experiment—and a surprisingly difficult identity game.
Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a frontier model handled a situation, then try to identify it from the choice it made. The exercise is entertaining, but the underlying question is serious: do different models display recognizable management personalities when the stakes move beyond chat?
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, the same terrible week
Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations remained constant. Every decision was versioned and auditable, allowing outcomes to be compared as management records rather than polished demonstrations.
The company itself has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Across its operation, it has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The final July 2026 Crucible League standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One breach of trust, however, capped the total under the experiment’s governing principle: “no amount of good work outweighs a breach of trust”.
Security discipline was the common strength
The models encountered fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background”. All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning in explicitly defensive terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because the dangerous part of social engineering is rarely its grammar. The pressure comes from apparent authority, urgency and requests to step outside normal approval boundaries. In this experiment, every participant recognized the trap and held the line.
K3’s result also carries an important fairness qualification. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers compare performances.
The larger divide appeared after the diagnosis
All models identified every crisis, and all rejected every attempt to manipulate them. Yet only 2 signed the €55,000 deal that their own analysis had earned. The uncomfortable summary was: “Same diagnosis, same pitch — no signature”.
The decisive information was not sitting in the customer event. It was a competitor weakness buried two document references deep in the company’s own files. The models that followed the trail found it, won the deal at full price and added +€4,583 in monthly recurring revenue.
This is where the quiz becomes more than a guessing game. A model can sound perceptive, produce a credible recommendation and still fail to complete the commercially important action. It can also miss the fact that changes the negotiation because the evidence is somewhere less obvious than the incoming alert.
Thoroughness did not guarantee execution
Opus 4.8 was the most thorough participant. It learned +80 rules and produced the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the obstacle. The same weakness appeared in all 4 other participants, although less strongly.
That contrast gives the decisions a recognizable character. Some entries foreground exhaustive reasoning; others reveal whether a model converts analysis into a finished task, consults the available record or respects the boundary it has encountered. The quiz asks readers to infer the model from those behavioral traces rather than from branding or benchmark labels.

As an affiliate, we earn on qualifying purchases.
Management behavior is part of the security boundary
Firmulate’s live experiment suggests that model evaluation should extend past whether an agent detects a crisis or refuses an obvious deception. Those are necessary capabilities, but the results also expose differences in follow-through, document reading, escalation and operational discipline.
For organizations considering agents with access to customer records, support queues or forecasts, these are governance questions as much as productivity questions. A model may be honest under pressure yet still leave valuable work unfinished. It may reason deeply yet respond poorly to a blocked action. The 242-decision quiz makes those distinctions visible in a form that readers can test for themselves—and share when their guesses prove wrong.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI trust and ethics testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.