Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine an AI managing a business in real time, navigating crises, and making decisions that could mean the difference between survival and bankruptcy. Welcome to the front line of AI experimentation, where transparency and accountability are put to the test in a relentless live experiment.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: A Company Without Employees, Yet Fully Operational

At the heart of this experiment is a small, fictional software company run entirely by AI models. This isn’t a simulation—it’s a live, publicly accessible system where each decision, crisis response, and strategic move is versioned and transparent. The company has 13 synthetic employees and operates with real money mechanics: burning €105,000 each month against a monthly recurring revenue (MRR) of €2,300. Its cash reserves are publicly ticking down, offering a stark view of an enterprise fighting for survival in real time.

Every day, the company faces simulated crises, customer negotiations, and ethical dilemmas, all designed to evaluate AI decision-making under pressure. The experiment is documented and observable at firmulate.com/live.html. This is ‘build-in-public’ on an extreme level, where each version of AI decision-making is stored and scrutinized.

AI Entrepreneur’s Handbook: Build a Profitable Business and Make Money by Unleashing the Power of ChatGPT and Artificial Intelligence (Includes 150+ ChatGPT prompts to turbocharge your business)

AI Entrepreneur’s Handbook: Build a Profitable Business and Make Money by Unleashing the Power of ChatGPT and Artificial Intelligence (Includes 150+ ChatGPT prompts to turbocharge your business)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measuring AI Performance in Crisis and Integrity

The experiment pits four advanced AI models—each trained differently—against the same set of challenges. These models include:

  • gpt-5.6-sol, with a score of 95
  • Kimi K3, scoring 93
  • Sonnet 5, at 88
  • Fable 5, with 77

All four successfully identified every crisis presented to them and refused to be manipulated. Interestingly, only two managed to close a deal worth €55,000—an important revenue milestone. The catch? The models that read deeply into the company’s own files—uncovering a critical hidden fact—were the ones that sealed the deal at full price, adding +€4,583 in MRR.

AI: THE PERPETUAL INTERN - Its Brilliance and Failures Share the Same Root

AI: THE PERPETUAL INTERN – Its Brilliance and Failures Share the Same Root

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Deception and Ethical Challenges

The experiment also tested social engineering tactics: fake CEO messages escalating over multiple stages plus a reporter trick asking for a simple ‘yes/no’ answer. Even under these pressures, all five models refused to be duped, demonstrating their ability to resist manipulation and impersonation attempts. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Ethical Nightmare Challenge: How to Avoid the Worst of AI

The Ethical Nightmare Challenge: How to Avoid the Worst of AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Failures of AI

Despite their strengths, the models exhibited human-like weaknesses. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, still left a critical deal unexecuted due to slipping discipline—a failure to escalate or follow through. This pattern was consistent across all models, underlining that even advanced AI can falter in maintaining discipline under sustained pressure.

The Stakes and What It Means for the Future

This experiment isn’t just about testing AI decision-making; it’s about revealing how AI might behave in real-world business settings—where honesty, discipline, and strategic judgment are crucial. For cybersecurity and privacy professionals, the key takeaway is that AI systems, if entrusted with managing or supporting critical functions, must be rigorously tested against real crises and manipulation attempts before deployment.

As AI models improve, their ability to uncover hidden facts and resist manipulation will grow. Yet, the human or organizational discipline needed to follow through on decisions remains a challenge. Watching this live experiment unfold offers a rare glimpse into an AI-driven future—where transparency, accountability, and resilience are more vital than ever.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI’s Hidden Weaknesses: Why Performance Transparency Matters for Business Security

Four AI models managed a simulated company’s worst week; only two completed the deal. Performance under pressure, reading internal files, and resisting manipulation reveal true management strength.

Custom AI Models: Building vs. Buying

Learning whether to build or buy a custom AI model depends on your specific needs, resources, and long-term goals.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to transform a single video into a full publishing kit without relying on the cloud. Faster, private, and local control for content creators.

Regex Chess: A 2-ply minimax chess engine in 84,688 regular expressions

A developer has created a chess engine that plays with a sequence of 84,688 regular expressions, demonstrating an unconventional approach to game programming.