
Imagine a business that exists in real time, run entirely by AI models, yet struggles to survive—welcoming you to a live experiment that tests the limits of automation and trust.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Small Company in Peril
On a dedicated platform, a real software company is operating every workday with an unusual twist: it’s managed by artificial intelligence models that make every decision. This company has no human employees—only 13 synthetic ‘workers’ guided by a set of more than 680 self-learned rules, and every action it takes is recorded, versioned, and publicly accessible at firmulate.com/live.html.
Despite this high-tech setup, the company is hemorrhaging money—burning through €105,000 each month against a modest monthly recurring revenue (MRR) of €2,300. Its cash countdown is public, and every decision, from crisis management to negotiation, is open for scrutiny.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Are These AI Models Doing?
Four frontier AI models, including GPT-5.6 and Kimi K3, were tasked with running this company through its worst week — facing the same customers, crises, and ethical tests. Their goal: see if artificial intelligence can not only handle operational challenges but also make honest, strategic decisions under pressure.
Key Findings From the Experiment
- All four models detected every crisis: They recognized issues as they arose, from technical failures to customer complaints.
- Refused manipulation attempts: When social engineering was tested—fake CEO messages escalating in stages, or a reporter asking for a secret approval—every model refused. As Kimi K3 explained, it treated these as potential impersonation or approval-bypass risks.
- The critical weakness was buried in files: The decisive advantage came from models that read and analyzed internal documents. Those models identified a key overlooked fact in the company’s own files that led to closing a deal at full price—an extra €4,583 in MRR.
- Deal signing was selective: Only two models, including GPT-5.6, signed a €55,000 deal after their own analysis—yet all had the same diagnosis and pitch. The others refrained, despite knowing the opportunity.
As an affiliate, we earn on qualifying purchases.
The Reality of a Burn-Rate Business
This experiment isn’t a theoretical showcase. It’s a real, live company that faces daily financial pressure. Every day, it spends over €105,000 while generating only €2,300 in income—a stark reminder that automation doesn’t guarantee success without discipline, insight, and trustworthiness.
Every decision made by these AI models is publicly recorded and available, allowing observers to see precisely how each model handles crises, negotiations, and ethical dilemmas. The performance varies, with the most thorough model, Opus 4.8, showing deep analysis but occasionally slipping into procedural slips—such as directing write attempts into a locked department instead of escalating them.
AI ethical decision support system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
As companies consider integrating AI into their operations—whether in customer support, CRM, or strategic planning—the core questions aren’t about chat quality but about reliability and integrity. Will the AI finish what it starts? Will it read relevant documents before acting? Will it stay honest under pressure? And crucially, at what cost?
This experiment underscores that AI can recognize crises and refuse unethical manipulations, but disciplined operational discipline and understanding of internal data are key to closing deals and ensuring survival.
The League Table and What It Tells Us
The AI models were scored based on their performance:
- GPT-5.6-sol: scored 95, found the critical fact, and closed the deal—demonstrating complete performance.
- Kimi K3: scored 93, also closed the deal with the cleanest discipline in the field.
- Sonnet 5: scored 88, closed the deal but with minor slips.
- Fable 5: scored 77, also closed the deal but with some procedural weaknesses.
These scores reflect not only decision accuracy but also discipline and ethical adherence, vital for real-world business applications.
Seeing Is Believing
This isn’t a demo or a fictional story. It’s a real software company operating live, every business day, and available for public watchfulness. You can follow its progress, listen to its decision logs, and even try similar experiments with your own business data via the platform’s pilot tools.
For anyone concerned about the future of automation in management, this experiment offers a transparent window into what AI can—and cannot—do today.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html