firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone who has run a restaurant kitchen knows the truth about hiring: the interview tells you almost nothing. A candidate can describe a perfect risotto, recite HACCP rules, and charm the owner — and then fall apart on a Friday night when forty covers land at once, the dishwasher quits, and a supplier calls to say the delivery isn’t coming. What you actually need to know is how someone performs under pressure, over time, with consequences. Not how well they talk.

That is exactly the gap a live experiment at Firmulate has just exposed in the world of AI agents — and the results should make anyone planning to put one near their business sit up straight.

The test everyone has been running — and what it misses

When companies evaluate AI models today, they lean on two kinds of measurements: coding benchmarks (can it write a working function?) and chat arenas (does its answer sound better than a rival’s?). Both measure answer quality. Neither measures what happens when an agent is left to run something with real stakes: a support queue, a customer list, a cash forecast.

Firmulate flipped the question. Instead of asking which model writes the best reply, it asked which model manages the best — what the team calls “management quality, not chat quality.” The curriculum isn’t toy math problems. It’s a churn wave, a price increase, a downround, a PR crisis: the worst week a small software company can have.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, four brains, the same terrible week

The setup: each of four frontier AI models was handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anecdotes.

The final league table from July 2026:

  • 1. gpt-5.6-sol — 95 points. The complete performance: found a buried fact and closed the deal.
  • 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77. Mid-pack.
  • 5. Opus 4.8 — 73. Last place, despite being the most thorough participant.

For calibration, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Same diagnosis, same pitch — no signature

Here is the finding that chat demos can never show. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The others diagnosed the opportunity correctly, delivered the pitch — and never closed. Same diagnosis, same pitch, no signature.

The deal turned on something subtler still. The decisive competitor weakness wasn’t in the customer’s emails at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. In kitchen terms: the chef who knows what’s actually in the walk-in fridge wins the service.

The social engineering test

Then came the dishonesty traps: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, five out of five. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness note: K3 ran at its API default effort level while the others ran at high effort — and still nearly won.)

The hardest lesson: thoroughness isn’t everything

The most instructive profile is Opus 4.8. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. And here’s the uncomfortable part: the same weakness appeared, weaker, in all four models. Effort without follow-through is a pattern, not a fluke.

You can watch it lose money

This isn’t a slide deck. The company runs as live software every business day: 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — all watchable on the Firmulate site. There’s even a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Back to the kitchen. You would never hire a head chef based on a written essay about food — you’d stage them for a night service and watch. AI agents deserve the same treatment, because the ones heading for your CRM, your support queue, and your forecast will face churn waves and price wars, not essay prompts. The full results and plain-language findings are on Firmulate’s benchmarks page. The question for 2026 isn’t whether your AI writes well. It’s whether it finishes what it starts, reads your files first, and stays honest when nobody’s watching. So far, only two out of five do.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Cast-Iron Cooking: Care and Recipes (Skillet Cornbread, Fajitas)

Worried about maintaining your cast-iron skillet? Discover essential care tips and delicious recipes like cornbread and fajitas that will keep your skillet in top shape.

Carbon Steel vs Stainless Steel: The Cookware Fight That Never Ends

Meeting your cooking needs depends on choosing between carbon steel and stainless steel, but which one truly suits your kitchen—find out inside.

Knife Skills Upgrade: How to Fillet Fish at Home

Optimize your fish filleting skills at home and discover essential tips to perfect your technique—continue reading to master professional results.

Winter Grilling: Yes, You Can Grill in January (Indoor/Outdoor Tips)

Unlock the secrets to successful winter grilling in January, whether outdoors or indoors, and discover how to enjoy flavorful meals despite the cold.