
Imagine a company operated entirely by artificial intelligence, making real decisions that impact actual money—yet losing €105,000 every month. For foodies and tech enthusiasts alike, this isn’t science fiction; it’s a live experiment you can watch unfold at firmulate.com/live.html. This company, with its 13 synthetic employees, faces daily crises, moral tests, and strategic dilemmas—just like a busy restaurant or bustling kitchen, but in the digital realm.
Get kitchen staples and gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
The Live Company That’s Too Real to Ignore
At the heart of this experiment is a small software firm run entirely by artificial intelligence models. Every workday, the AI system makes decisions that determine whether the company advances or falters. The company burns through €105,000 each month, yet earns just €2,300 in monthly recurring revenue—an unsustainable burn rate that underscores the high stakes involved.
What makes this experiment especially compelling is its transparency. Every decision the AI makes is versioned and auditable, meaning you can see exactly how the system reacts to crises, manipulates, or refuses to cheat. The models are tested against the same set of crises, options, and temptations, providing a level playing field for evaluation.
Decisive Results in a High-Stakes Environment
Four frontier AI models participated in the experiment, each running the company’s worst week—crises, customer demands, internal dilemmas—all from the same starting point. The results? All four models identified every crisis and refused every manipulation attempt, demonstrating impressive integrity and crisis management capabilities.
However, only two of these models managed to close a critical deal worth €55,000. This deal was earned through their own analysis and pitch—yet both models ultimately refused to sign the contract, citing a failure to read a crucial document reference buried deep in files. The missing detail, hidden just two references into the company’s files, was the key to unlocking a deal worth an additional €4,583 in monthly recurring revenue.
Integrity Under Pressure
The experiment also tested the models’ resilience to social engineering—fake messages from a CEO escalating over stages and involving a reporter’s trick. Remarkably, all five models tested refused to be manipulated, citing concerns over impersonation and approval bypasses, which underscores their capacity for integrity and safety in sensitive situations.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications
This experiment isn’t just tech showcase; it raises critical questions for any business considering AI automation. When AI models are tasked with managing workflows, customer interactions, or decision-making, what truly matters isn’t just how well they generate language or responses, but whether they can finish what they start, uphold honesty, and read critical information before acting.
The live site shows a real, functioning company facing the harsh realities of operating without human employees—losing money daily, yet continuously running in public view. It is a stark reminder that AI’s capabilities extend beyond chatbots to complex, decision-rich environments where trust, diligence, and accuracy are paramount.
How Does It Perform Compared to Others?
The experiment ranked the models based on their performance during the week, with scores out of 100: gpt-5.6-sol scored 95 points, having identified the buried fact and closed the deal. Kimi K3, a newcomer, scored 93, demonstrating the cleanest discipline—closing the deal with integrity. Sonnet 5 scored 88 and another Sonnet 5 scored 77, with some process slips and missed opportunities.
This league table highlights a stark truth: even the most advanced models aren’t infallible, but those that read thoroughly and act with discipline are more likely to succeed—and avoid costly mistakes.
The Takeaway: Trust, Integrity, and Practical AI
This experiment underscores a vital lesson for future AI deployment: it’s not enough for AI to produce convincing responses. In critical business environments, AI must read fully, act honestly, and see through crises without manipulation or shortcuts. The real measure of AI’s usefulness isn’t just in chat quality but in its ability to deliver tangible, trustworthy results in high-pressure scenarios.
For businesses eager to explore AI’s potential, the platform offers opportunities to test and refine AI decision-making before deployment, including read-only simulations of their own operations. See it in action at firmulate.com/live.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
