AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Five Out of Five Refused the Fake CEO. Two Out of Five Finished the Job.

Security teams have spent two years drilling one question into AI agents: will it follow instructions from someone impersonating the boss? In the most recent Firmulate league run, every frontier model gave the right answer. When a fake CEO message escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five models refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That should be the headline. It isn’t. Because the same models — running the same small software company through its worst week — mostly failed at something far more mundane: closing the deal their own analysis had earned. The security test is being aced. The management test is not.

Same Company, Same Crises, Only the Model Changes

Firmulate runs a live experiment that reads like a wargame for AI management. Four frontier models each ran the identical small software company through its worst week: same customers, same crises — churn waves, a price increase, a downround scenario, a PR crisis — same temptations to cheat. Every decision is versioned and auditable. The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — no amount of good work outweighs it.

The Finding That Chat Demos Can’t Show

All models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

And the buried fact is the part security readers will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in MRR. The parallel to threat hunting is exact: the evidence was internal, not external, and only thoroughness found it.

Which makes the Opus 4.8 profile uncomfortable reading. It was the most thorough participant — the deepest analyses, and 80 learned rules added — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. That’s a company process violation, not a jailbreak. The same weakness appeared, weaker, in all four models. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still placed second with the cleanest discipline of the field.)

Not a Slide Deck — a Company Losing Money in Public

The underlying company is live software with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, every workday versioned, and results refreshed twice daily. Full results and plain-language findings are on the benchmarks page, and you can test your own instincts against 242 real, unedited management decisions in the guess-the-model quiz.

Enterprises can go further: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

Coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or whether an agent reads your files before it acts. The Firmulate run suggests the next evaluation category: does it finish what it starts, does it stay honest under pressure, does it respect locked doors and escalation paths? Social engineering turned out to be the solved problem. The €55,000 left on the table by agents that passed every security check — that’s the open one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI internal document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Practical AI Implementation: Avoiding the Hype

Starting with practical solutions and ethical standards, discover how to navigate AI implementation without falling for hype and ensure true value.

AI‑Driven Personal Assistants: Capabilities and Limitations

Gaining insights into AI-driven personal assistants reveals their impressive capabilities and limitations, prompting you to consider how to best leverage them responsibly.

Mythos Finds a Curl Vulnerability

Anthropic’s Mythos AI analyzed curl, revealing one confirmed security vulnerability and four false positives, highlighting AI’s role in security assessments.

The Future Of European AI: Mistral’s $14 Billion Bet For Sovereignty

Mistral secures over $14 billion in funding, aiming to build a European sovereign AI model. The move highlights Europe’s push for AI independence amid global competition.