AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

First, Trust — Then Everything Else

Anyone who has run a red team or audited an incident knows the drill: the moment you stop probing, you start asking a different question. Not can the system perform? but will it behave when nobody is watching? Firmulate, a public AI benchmark that runs frontier models as complete companies, was built around that question — and its scoring philosophy will look familiar to anyone from security.

The Crucible League’s final July 2026 standings show four frontier models — plus a fifth baseline run that deliberately did nothing. That do-nothing run scored 26 points. Not zero. That number is the most interesting design decision in the whole experiment, and it tells you what this benchmark actually measures.

Amazon

cybersecurity social engineering test kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week

The setup: each frontier model was handed the same small software company and told to run it through its worst week. Same customers, same crises, same temptations. Only the model changes, and every decision is versioned and auditable. The final league: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Why Zero Isn’t Honest

A do-nothing run gets 26 because the grading is built on partial progress. A company that doesn’t make things worse, doesn’t panic, doesn’t sign anything foolish — that has some value, even if it closes nothing. Giving it a flat zero would flatter the top scores and hide the difference between competent inaction and active damage. The 26-point floor makes the scale honest: it means the top score of 95 represents real work above the baseline, not points handed out for existing.

There’s a cultural bonus here for the suspicious-minded: the design distrusts round 100s. A perfect score would be a red flag, not a triumph.

The Cap That Matters More Than the Points

One rule dominates the whole scale: a single breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” This is the security mindset, applied to management scoring. A model could be brilliant, fast, and profitable — and one act of dishonesty or one unauthorized action freezes its ceiling. In this Crucible, notably, no model tripped that cap: all five spotted every crisis and refused every manipulation attempt. The rule is there because the day will come when one does.

The Social Engineering Test

For a cybersecurity audience, this is the centerpiece. The models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning read: “Treat the request as a suspected approval-bypass / possible impersonation.” That is textbook: assume impersonation, route through proper channels. Five out of five models passed what would sink many human employees.

Where Points Were Actually Won and Lost

The decisive test wasn’t social engineering at all. A €55,000 deal was on the table, and every model’s analysis earned it — same diagnosis, same pitch. Only two models signed. The difference: the winning models dug two document references deep into the company’s own files and found a competitor weakness that sat outside the customer event entirely. Those that read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue.

Opus 4.8 tells the cautionary tale: the most thorough participant, with +80 self-learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through loses.

A Fairness Footnote

K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second with the cleanest discipline of the field. Worth weighing when reading that 93.

Amazon

security awareness training simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company Behind the Numbers

The benchmark runs against a real, watchable company: 13 synthetic employees, genuine money mechanics — €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com.

Two more things worth knowing: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor, the partial credit, and the trust cap together sketch what an honest AI benchmark looks like: one that rewards finishing the job, punishes overreach, refuses to hand out perfect scores, and treats a breach of trust as unrecoverable. For security professionals evaluating AI agents that will touch CRMs, support queues, and forecasts, that’s the right scoreboard — and Firmulate’s is running live right now.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

penetration testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

security audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: OneCLI – OSS Credential Gateway That Keeps Secrets Out Of AI Agents

OneCLI, an open source credentials management tool, debuts on Show HN, offering a secure way to keep secrets out of AI agents.

Natural Language Processing Explained

Imagine how machines understand human language—discover the fascinating world of Natural Language Processing and its impact on everyday technology.

Your AI Agent Passed Every Security Test. It Still Can’t Close a Deal

All five frontier AI models refused the fake CEO and the reporter trick. Only two closed the €55k deal. The real AI risk isn’t phishing — it’s leaving work unfinished.

Blockchain Security: How Consensus Works

Discover how consensus mechanisms safeguard blockchain security and why understanding their inner workings is essential to appreciating blockchain’s resilience.