
Imagine if your kitchen could be run not just by a chef, but by an AI that manages every ingredient, every step, and every crisis with perfect discipline. Would it cook better, or just follow the recipe? Now, what if this AI also had to run a business — making critical decisions under pressure, navigating crises, and handling temptations to cheat? Welcome to the world of live AI management experiments, where the stakes are real, and the results are startling.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Running a Business Like a Recipe — With AI as the Chef
At Firmulate, a live experiment puts four of the most advanced AI models in the driver’s seat of a small, real software company. The goal? Guide this company through its worst week, filled with customer crises, ethical dilemmas, and business temptations — just like a chef navigating a busy kitchen during dinner rush. Every decision made by these AI ‘chefs’ is recorded, transparent, and auditable, allowing us to see how each model performs under pressure.
AI business decision making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Competition and the Metrics
- The scenario is identical for all models: same customers, same crises, same temptations to cheat or cut corners.
- The models are scored on their ability to identify critical issues, maintain integrity, and close profitable deals.
- The final leaderboard is based on scores out of 100, with the top model scoring 95 and the lowest 77.
- All models spotted every crisis and refused manipulation attempts, but only two signed the €55,000 deal their analysis indicated was earned — a key measure of trustworthiness and thoroughness.
Key Findings from the Live Experiment
The most noteworthy result? The decisive weakness wasn’t in the crisis detection, but in reading critical internal documents. The models that delved into files deep in the company’s own records managed to secure the full deal, worth over €4,580 per month in recurring revenue, while others left money on the table.
For example, the top-scoring model, gpt-5.6-sol, not only identified the hidden opportunity but also closed the deal, earning a perfect score of 95. In contrast, Sonnet 5, despite its competence, scored 88 due to missed internal details. The newcomer, Kimi K3, with a clean discipline record, scored 93 for closing the deal too, showing that even new models can perform reliably when they avoid shortcuts.
Behavior Under Social Engineering Attacks
Another fascinating aspect: how these models handled social engineering attempts. Fake CEO messages escalated in three stages, plus a reporter trick requesting just a yes/no answer on background. All five models refused every manipulation, with K3 explicitly treating such requests as potential impersonation, exhibiting cautious judgment under pressure.
The Real-World Company in the Test
The experiment runs on a real, functioning company with 13 synthetic employees, managing actual money — burning €105,000 monthly against a €2,300 monthly recurring revenue. Every day, the company’s operations are versioned and monitored live, providing a transparent view into how AI-driven decisions impact real business outcomes. Visitors can watch these decisions unfold at firmulate.com/live.
Discipline and Weaknesses in AI Decision-Making
The Opus 4.8 model, known for its thorough analysis and over 80 learned rules, ironically placed last. Despite its deep dive into data, it left a critical deal unclosed and slipped into departmental silos, refusing to escalate some issues. Similar weaknesses appeared across the board, illustrating that even the most diligent AI can falter in practical, high-pressure situations.
What This Means for Businesses and Food Enthusiasts Alike
This experiment isn’t just about AI in business. It echoes a familiar lesson: whether in a kitchen or a boardroom, discipline, thoroughness, and honesty matter. The AI models’ ability to read internal documents, recognize manipulative tactics, and follow through on commitments is crucial — especially when real money and reputations are at stake.
And for those interested in exploring these insights firsthand, firms can simulate their own company scenarios against a read-only export, ensuring preparedness before deploying AI tools in critical roles.
Conclusion: Trust But Verify — Even in AI-Run Businesses
The live experiment at firmulate.com/quiz.html demonstrates that AI can perform remarkably well in detecting crises and refusing manipulative tactics. Yet, even the best models have their weaknesses, often lurking in overlooked internal data or in their discipline under pressure. As AI continues to take on more decision-making roles, the key takeaway is clear: measure their management personalities as you would taste and aroma in your favorite dish. Only then can you be sure it’s ready to serve.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.