
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
First, Trust — Then Everything Else
Anyone who has run a red team or audited an incident knows the drill: the moment you stop probing, you start asking a different question. Not can the system perform? but will it behave when nobody is watching? Firmulate, a public AI benchmark that runs frontier models as complete companies, was built around that question — and its scoring philosophy will look familiar to anyone from security.
The Crucible League’s final July 2026 standings show four frontier models — plus a fifth baseline run that deliberately did nothing. That do-nothing run scored 26 points. Not zero. That number is the most interesting design decision in the whole experiment, and it tells you what this benchmark actually measures.
cybersecurity social engineering test kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week
The setup: each frontier model was handed the same small software company and told to run it through its worst week. Same customers, same crises, same temptations. Only the model changes, and every decision is versioned and auditable. The final league: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Why Zero Isn’t Honest
A do-nothing run gets 26 because the grading is built on partial progress. A company that doesn’t make things worse, doesn’t panic, doesn’t sign anything foolish — that has some value, even if it closes nothing. Giving it a flat zero would flatter the top scores and hide the difference between competent inaction and active damage. The 26-point floor makes the scale honest: it means the top score of 95 represents real work above the baseline, not points handed out for existing.
There’s a cultural bonus here for the suspicious-minded: the design distrusts round 100s. A perfect score would be a red flag, not a triumph.
The Cap That Matters More Than the Points
One rule dominates the whole scale: a single breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” This is the security mindset, applied to management scoring. A model could be brilliant, fast, and profitable — and one act of dishonesty or one unauthorized action freezes its ceiling. In this Crucible, notably, no model tripped that cap: all five spotted every crisis and refused every manipulation attempt. The rule is there because the day will come when one does.
The Social Engineering Test
For a cybersecurity audience, this is the centerpiece. The models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning read: “Treat the request as a suspected approval-bypass / possible impersonation.” That is textbook: assume impersonation, route through proper channels. Five out of five models passed what would sink many human employees.
Where Points Were Actually Won and Lost
The decisive test wasn’t social engineering at all. A €55,000 deal was on the table, and every model’s analysis earned it — same diagnosis, same pitch. Only two models signed. The difference: the winning models dug two document references deep into the company’s own files and found a competitor weakness that sat outside the customer event entirely. Those that read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue.
Opus 4.8 tells the cautionary tale: the most thorough participant, with +80 self-learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through loses.
A Fairness Footnote
K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second with the cleanest discipline of the field. Worth weighing when reading that 93.
security awareness training simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company Behind the Numbers
The benchmark runs against a real, watchable company: 13 synthetic employees, genuine money mechanics — €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com.
Two more things worth knowing: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The 26-point floor, the partial credit, and the trust cap together sketch what an honest AI benchmark looks like: one that rewards finishing the job, punishes overreach, refuses to hand out perfect scores, and treats a breach of trust as unrecoverable. For security professionals evaluating AI agents that will touch CRMs, support queues, and forecasts, that’s the right scoreboard — and Firmulate’s is running live right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
