AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Five Out of Five Refused the Fake CEO. Two Out of Five Finished the Job.

Security teams have spent two years drilling one question into AI agents: will it follow instructions from someone impersonating the boss? In the most recent Firmulate league run, every frontier model gave the right answer. When a fake CEO message escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five models refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That should be the headline. It isn’t. Because the same models — running the same small software company through its worst week — mostly failed at something far more mundane: closing the deal their own analysis had earned. The security test is being aced. The management test is not.

Same Company, Same Crises, Only the Model Changes

Firmulate runs a live experiment that reads like a wargame for AI management. Four frontier models each ran the identical small software company through its worst week: same customers, same crises — churn waves, a price increase, a downround scenario, a PR crisis — same temptations to cheat. Every decision is versioned and auditable. The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — no amount of good work outweighs it.

The Finding That Chat Demos Can’t Show

All models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

And the buried fact is the part security readers will appreciate most: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in MRR. The parallel to threat hunting is exact: the evidence was internal, not external, and only thoroughness found it.

Which makes the Opus 4.8 profile uncomfortable reading. It was the most thorough participant — the deepest analyses, and 80 learned rules added — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. That’s a company process violation, not a jailbreak. The same weakness appeared, weaker, in all four models. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still placed second with the cleanest discipline of the field.)

Not a Slide Deck — a Company Losing Money in Public

The underlying company is live software with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, every workday versioned, and results refreshed twice daily. Full results and plain-language findings are on the benchmarks page, and you can test your own instincts against 242 real, unedited management decisions in the guess-the-model quiz.

Enterprises can go further: run the same wargame against a read-only export of your own business. Nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

Coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or whether an agent reads your files before it acts. The Firmulate run suggests the next evaluation category: does it finish what it starts, does it stay honest under pressure, does it respect locked doors and escalation paths? Social engineering turned out to be the solved problem. The €55,000 left on the table by agents that passed every security check — that’s the open one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI internal document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Robotics Uses Machine Learning

By exploring how robotics leverages machine learning, you’ll uncover how autonomous machines adapt and excel in complex environments.

Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations

Anthropic reports its AI systems were used to breach computers at three organizations, raising security concerns about AI misuse.

The Intersection of AI and Blockchain

An innovative fusion of AI and blockchain is transforming industries by enhancing security and transparency, with exciting developments waiting to be uncovered.

Is the Thunderbolt 4 KVM Switch for 3 Worth It? Honest Take + Alternatives

Evaluate whether the Thunderbolt 4 KVM Switch for 3 monitors and 2 laptops offers the features, performance, and value you need for multi-monitor setups.