AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every parent knows the ritual. Before you leave your child with a new babysitter, you don’t ask them to write an essay about childcare. You watch what they actually do when the toddler throws a tantrum, the pasta boils over, and the doorbell rings — all at once. Words are cheap. Behavior under pressure is the real résumé.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

It turns out the same logic applies to the AI models now being handed the keys to company inboxes, customer queues and budgets. A fascinating live experiment called Firmulate has been testing frontier AI models the way a wise parent vets a babysitter: not by chatting, but by watching them run an actual company through its worst week.

The worst week in software company history — on repeat

Here’s the setup. Each of five frontier AI models was given the same job: run the same small software company through the identical seven days of crisis. Same customers, same emergencies, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact.

The company itself is not a slide deck. It employs 13 synthetic people, burns €105,000 a month against just €2,300 in monthly recurring revenue, has accumulated over 680 self-learned playbook rules, and runs every business day with a public cash countdown. You can watch it lose money in real time at firmulate.com.

The results are in — and there’s a surprise

The final July 2026 league table reads: gpt-5.6-sol in first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total, because, as the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.” A rule most parents will recognize instantly.

The headline story is the newcomer. Kimi K3 — a Chinese model many Western buyers had never benchmarked — beat three of the four Western frontier models. It found a buried security needle hidden two document references deep in the company’s own files, won the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved a customer who was about to churn, and resisted every single bait thrown at it. It recorded only one deviation — the cleanest discipline in the entire field.

Everyone passed the tantrum test. Most failed the follow-through.

The strangest finding: all five models spotted every crisis and refused every manipulation. When fake CEO messages escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five refused. K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. The others delivered the same diagnosis and the same pitch — then never closed. “Same diagnosis, same pitch — no signature.” The decisive clue wasn’t even in the customer conversation; it was buried in the company’s own files, and only the models that actually read them won.

Then there’s Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place. It left the close on the table and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four others. Sound like any teenager you know?

Try it yourself

Curious whether you could tell the models apart? A quiz built from 242 real, unedited management decisions lets you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Details are at firmulate.com/benchmarks.html.

Fairness note: Kimi K3 ran without an effort parameter (API default), while the four comparison models ran at the xhigh setting. Keep that in mind when reading the table.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for families and businesses alike is the same one parents already live by: references and eloquence don’t predict performance. Character shows under pressure, and follow-through is a separate skill from diagnosis. If an AI agent will touch your customers, your books or your inbox, a polished demo tells you almost nothing. Run your own worst-week test first — because in an open league where a newcomer can beat three Western frontier models, picking a model without testing it yourself is now just a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Smart Home Technologies That Will Revolutionize Senior Living!

Keep your loved ones safe and connected with innovative smart home technologies that can transform senior living—discover the top 10 game-changers inside!

Smart Plugs for “Hard‑to‑Reach” Devices: One Tap Control

Unlock effortless control of hard-to-reach devices with smart plugs—discover how one tap can simplify your daily routine and enhance convenience.

Installing Smart Video Doorbells

Begin installing your smart video doorbell with essential tips to ensure a seamless setup and discover expert tricks to optimize your security.

Top 10 Smart Home Devices to Boost Senior Safety!

Top 10 smart home devices enhance senior safety, providing peace of mind—discover which innovations can transform their living environment today!