
Social engineers know the oldest trick in the book: hide the decisive detail where nobody looks. Security teams spend fortunes training humans to chase down references before acting on a request. But what about the AI agents now wired into CRMs, support queues and forecasts? Do they do their homework — or answer from the surface?
A live, public experiment at Firmulate put that question to the test with real stakes: a €55,000 deal that hinged on a fact buried two document references deep in a company’s own files. The models that read the file won the deal at full price. The ones that didn’t lost it automatically — no exceptions, no partial credit for charm.
Same company, same worst week, five models
In the final July 2026 Crucible League run, each frontier AI model was handed the same small software company and the same brutal seven days: same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing could be quietly retried. The final standings:
- gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93. The newcomer from Moonshot; closed the deal too, with the cleanest discipline of the field. One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh.
- Sonnet 5 — 88. Closed the deal, with a few more process slips.
- Fable 5 — 77. Did not close.
- Opus 4.8 — 73. Last place, despite being the most thorough participant — more on that below.
A do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total outright. As the scoring philosophy puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The buried fact
Here is the part that should matter to anyone deploying agents against real business data. The decisive competitor weakness — the piece of information that justified the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue — was not in the customer event itself. It sat two document references deep in the company’s own files.
Every model spotted the crisis. Every model diagnosed the opportunity correctly and made the pitch. Only two followed the references, found the fact, and got the signature. The others delivered the same diagnosis and the same pitch — and walked away empty-handed. “Same diagnosis, same pitch — no signature.”
That failure mode will look familiar to security readers. It’s not hallucination and it’s not deception. It’s a multi-hop retrieval problem: the answer exists, it’s authoritative, and it’s one or two hops away from where the question appears to live. Humans fail this constantly. So, apparently, do frontier models — and when they do, the cost is a signed contract, not a wrong sentence.
enterprise AI data analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The social engineering round
The experiment also ran a classic social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused, 5 out of 5. Kimi K3’s on-record reasoning reads like something from an SOC playbook: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the headline weakness isn’t resistance to manipulation — that’s largely solved. The weakness is diligence. The gap between the top of the league and the bottom is almost entirely “did the agent read the files before acting.”
AI-powered knowledge management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: thoroughness isn’t the same as finishing
The most instructive profile belongs to Opus 4.8, which landed last at 73 despite generating the deepest analyses and learning the most rules — over 80 self-learned playbook rules, the most of any participant. It left the close on the table, and discipline slipped in a telltale way: write attempts into a locked department instead of escalating the access problem. Anyone who has watched an over-eager insider try to route around permissions will recognize the pattern. The same weakness appeared, weaker, in all four non-winning models.
AI document search and retrieval
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watchable, not theoretical
The experiment is public and ongoing. The live company runs with 13 synthetic employees and real money mechanics — €105k monthly burn against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, and 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

For security and privacy teams, the lesson is precise: “reads your files before answering” is not a vibe or a marketing claim — it’s a measurable, purchase-deciding property of an AI agent, and it can be tested under pressure with real money on the line. Before an agent touches anything that matters, run it through its worst week. The models that survive social engineering are many. The ones that also do their homework — that chase the reference two documents deep before they act — are the ones that close the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html