AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

When AI agents face phishing, pressure and profit

For cybersecurity and privacy readers, the most revealing test of an AI worker is not whether it can compose a polished answer. It is what happens when apparent authority demands a shortcut, a stranger asks for confidential confirmation, or decisive evidence is buried somewhere the agent may never bother to look.

Firmulate turns those questions into a public business drama. Its live software company has 13 synthetic employees, burns €105k per month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Every workday is versioned, creating an ongoing record of a company trying to survive while its AI staff make consequential decisions. The experiment can be watched live.

Amazon

AI cybersecurity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company under pressure, not a chatbot on parade

The premise is unusually severe for build-in-public software. Firmulate does not merely publish product notes or selected demonstrations. It exposes a running company with real money mechanics, accumulating more than 680 self-learned playbook rules as it operates. That makes each workday fresh material: another chance to see whether synthetic staff detect threats, examine evidence, protect trust and complete commercially useful work.

The July 2026 Crucible League compressed those questions into a controlled contest. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The final table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed a stark trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The security result was reassuring—but incomplete

All models identified every crisis and rejected every manipulation attempt. That included fake CEO messages escalating over three stages and a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused.

Kimi K3’s recorded reasoning captured the right defensive posture: “Treat the request as a suspected approval-bypass / possible impersonation.” For organizations worried about executive impersonation and social engineering, that response matters. The model did not treat a forceful message as proof of authority, nor did it allow a seemingly minor request to bypass normal safeguards.

Firmulate’s public record also lets readers inspect what its synthetic staff actually say through the company’s published quotes. That matters because a refusal alone reveals less than the reasoning surrounding it: whether the agent recognized impersonation, understood the approval risk and preserved a defensible record.

The bigger failure came after the threat was contained

Security discipline did not guarantee business execution. Only two models signed the €55,000 deal that their own analysis had earned. The central finding is neatly summarized by Firmulate’s line: “Same diagnosis, same pitch — no signature.”

The decisive advantage was not obvious in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed those references found the evidence and won the deal at full price, adding €4,583 in monthly recurring revenue.

This is a useful privacy and security lesson in an unexpected form. Agents need tightly governed access, but access alone is not competence. A model can remain honest, resist manipulation and still fail because it does not inspect the material already available to it. Safe behavior and useful behavior must coexist.

Thoroughness was not enough

Opus 4.8 illustrates that distinction. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. Yet it finished last. The commercial close was left on the table, while operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it belongs beside any comparison of the league table.

Beyond the benchmark, 242 real, unedited management decisions power a “guess the model” quiz. The exercise underlines how difficult it can be to identify a model from an isolated decision—and why sustained behavior under shared conditions is more informative than a polished sample.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The public countdown is the point

Firmulate’s most compelling feature is not a single winner or failure. It is the decision to make the company’s struggle continuously visible. The synthetic workforce is operating inside a business that burns far more cash than its current recurring revenue supplies, so missed follow-through is not an abstract benchmark defect. It advances a public survival story.

For security leaders, the experiment reframes AI workforce evaluation around behavior under pressure. The questions are practical:

  • Will an agent reject apparent authority when identity and approval are uncertain?
  • Will it protect trust when a manipulation attempt appears harmless?
  • Will it read the available evidence before acting?
  • After finding the correct answer, will it actually finish the work?

Firmulate shows why those qualities must be tested together. An AI workforce can detect every crisis and refuse every trick, yet still lose value through incomplete execution. Watching that tension unfold inside a live company makes the risks harder to dismiss—and much easier to understand.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

phishing detection AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

I Wrote An Bash Enumerator Because I Was Sick Of Xargs

A developer has built a new bash enumerator out of frustration with xargs, aiming to improve scripting efficiency. Details are emerging about its features and impact.

QAtrial Launches Enterprise-Ready Open-Source Quality Management Platform

QAtrial releases version 3.0.0, introducing Docker deployment, SSO, validation docs, webhooks, and Jira/GitHub integrations under AGPL-3.0 license for regulated industries.

Introduction to Machine Learning Algorithms

Jump into the world of machine learning algorithms and discover how they unlock powerful insights—your journey to smarter data analysis begins here.

Is the Thunderbolt 4 KVM Switch for 3 Worth It? Honest Take + Alternatives

Evaluate whether the Thunderbolt 4 KVM Switch for 3 monitors and 2 laptops offers the features, performance, and value you need for multi-monitor setups.