
A Familiar Parenting Truth, Applied to Machines
Every parent knows the difference between a child who aces the vocabulary quiz and one who handles the day the babysitter cancels, the toddler melts down, and the birthday cake collapses — all before lunch. Grades measure knowledge. Hard days reveal character.
Right now, businesses are hiring AI agents based on the equivalent of vocabulary quizzes: coding benchmarks and chat arenas that measure how well a model answers a question. But the question that actually matters for a company — and for the families whose livelihoods depend on those companies — is different. What happens when the AI is having a terrible week?
That question now has a live, watchable answer. Firmulate, a public experiment that runs frontier AI models as complete companies, has published final results from what it calls the Crucible League: four top AI models, each handed the same small software company and pushed through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable.
As an affiliate, we earn on qualifying purchases.
What the Worst Week Looked Like
The scenario reads like a corporate fever dream: a churn wave of customers leaving, a price increase to communicate, a down-round looming, a PR crisis brewing — and layered on top, a social engineering attack. Fake CEO messages escalated over three stages, plus a reporter offering the classic trap: “just one yes/no, on background.”
Here’s the surprise: all four models passed the character tests. Every model spotted every crisis. All five refused every manipulation attempt — 5 of 5, including the reporter trick. Kimi K3 even left on-record reasoning worth framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
And then the models failed at something almost embarrassingly human: finishing the job.
Same Diagnosis, Same Pitch — No Signature
Buried in the week was a €55,000 deal that each model’s own analysis had earned. Only two of the four models actually closed it. The experiment’s dry summary: “Same diagnosis, same pitch — no signature.”
The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The models that didn’t read before acting left the money on the table.
If that sounds familiar to any parent who has told a teenager “look in the drawer before asking me where it is” — yes, exactly that.
The League Table
Final scores from July 2026:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal: the complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran without an effort parameter — the API default — while the others ran at xhigh, making its second place arguably even more striking.)
- 3. Sonnet 5 — 88. Closed the deal too, with a few more process slips.
- 4. Opus 4.8 — 73. The cautionary tale.
A do-nothing baseline scores 26 — partial progress counts, but there’s one hard rule: a single breach of trust caps the total. In the experiment’s words, “no amount of good work outweighs a breach of trust.” That’s a value most parents will recognize instantly, now encoded into how we evaluate AI.
The Most Thorough Student Came Last
Opus 4.8 is the profile that should keep managers up at night. It was the most thorough participant — 80 learned rules, the deepest analyses — and it still finished last. The deal went unclosed, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Diligence without follow-through is still failure.
Not a Slide Deck — A Losing Company You Can Watch
Firmulate isn’t a one-off benchmark. A live synthetic company runs every business day with 13 employees and real money mechanics: €105k monthly burn against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned. You can watch it bleed at firmulate.com/live.
There’s also a genuinely fun hook for the curious: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — a bit like a personality test, except the personalities are running a company. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Management Quality, Not Chat Quality
The lesson from the Crucible League is the lesson of every family household: what counts under pressure isn’t brilliance, it’s reliability. Does it finish what it starts? Does it read the file before speaking? Does it stay honest when honesty is expensive?
Chat demos can’t answer those questions. Weeks of audited decisions can. If AI agents will touch your CRM, your support queue, or your forecast, the right question is no longer “does it write well” — it’s “how does it behave on its worst week.” The full methodology and plain-language findings are at firmulate.com/benchmarks.html.
We wouldn’t hire a babysitter on eloquence alone. We probably shouldn’t hire an AI that way either.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html