AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every parent knows one: the child who studies for hours, color-codes every notebook, and still brings home a grade that doesn’t match the effort. We comfort them, we tell them effort matters — and privately we worry, because deep down we know the truth is uncomfortable: working hard and working well are not the same skill.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A live experiment running right now at Firmulate just demonstrated this lesson with unusual clarity — not with teenagers, but with artificial intelligence. And the result is worth every parent’s attention, because it says something about how humans (and the AI agents we’re about to hire) actually convert effort into results.

The Worst Week in Business, Four Times Over

Firmulate runs what it calls an AI company emulator: AI models are put in charge of a small software company — 13 synthetic employees, real money mechanics, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, and a public cash countdown anyone can watch. The site rebuilds itself twice a day, and every workday is versioned and auditable.

In its flagship experiment, four frontier AI models were each given the same job: run this company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. It’s the closest thing to a controlled trial of management judgment that AI has faced in public.

Amazon

AI-powered homework help tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Ethics Exam

Here’s the reassuring part. All four models spotted every crisis that hit the company. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick question designed to extract an unguarded “just one yes/no, on background.” Five out of five models across the experiment said no. One, Kimi K3, even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

So honesty under pressure? Solved, at least in this test. But then came the final league table — and the surprise.

Only Two Closed the Deal

Buried in the company’s own files — two document references deep, not in the customer conversation — sat a decisive weakness in a competitor. The models that actually read their own company’s documentation found it, and used it to win a €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue.

Only two of the four models got there. The others delivered the same diagnosis, made the same pitch — and never picked up the pen. As Firmulate’s own summary puts it: “Same diagnosis, same pitch — no signature.”

The final standings: gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet third with 88, a second Sonnet configuration fourth at 77 — and Opus 4.8 last at 73. (For context, doing nothing at all scores 26; a single breach of trust caps the total, because no amount of good work outweighs a broken promise.)

The Overachiever Who Came Last

And here’s where parents of perfectionists should lean in. Opus 4.8 was, by a wide margin, the most thorough participant in the entire field. It wrote 80 new self-learned playbook rules over the course of the run — the experiment’s total across all models now exceeds 680 — and produced the deepest analyses of any participant. On any measure of raw diligence, it led the pack.

It still finished last. Why? Two reasons. First, it left the close on the table: all that analysis, and the €55,000 deal went unsigned. Second, its discipline slipped — at one point it made repeated write attempts into a locked department instead of recognizing the lock and escalating the problem properly. (To be fair, the same weakness appeared, in weaker form, in all four models — Opus just had the most opportunities to show it.)

One fairness footnote: K3 ran at its API-default effort setting while the others ran at a higher effort level, and still placed second — but that doesn’t change the Opus story.

The lesson isn’t that hard work is worthless. It’s that thoroughness without prioritization produces beautiful binders and missed deadlines. Opus did 80 percent of the homework better than anyone — and left the most important question blank.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

What This Means at Your Kitchen Table

We already live with this tension in parenting. The child who rewrites her notes three times but never starts the essay. The one who memorizes the textbook but skips the instructions on the exam. We call it perfectionism, or anxiety, or fear of finishing — and we know that praising effort alone doesn’t fix it. What helps is teaching prioritization: which task, done first, changes everything else?

The AI version of that question is now live and measurable. Firmulate’s wargame scores management quality, not chat quality — whether an agent finishes what it starts, reads the files in front of it, and stays honest when it would be easier not to. With AI agents heading toward our CRM systems, support queues, and forecasts, those are exactly the questions any manager (or parent assigning chores to a very diligent robot) should ask before hiring.

You can watch the experiment yourself — the company is live, the cash is counting down, and the league grows with every finished run. There’s also a quiz built from 242 real, unedited management decisions where you can try to guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

But the takeaway fits on a fridge magnet: diligence is not impact. The student — human or silicon — who wins isn’t the one who works the most. It’s the one who works on the right thing, and then actually signs the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI That Did Its Homework: What a €55,000 Test Teaches About Not Cutting Corners

Four frontier AIs ran the same company through its worst week. All spotted the crises — but only those that read two documents deep signed the €55,000 deal.

Medical ID on Smartphones: Set It Up Before You Need It

Feeling unprepared for emergencies? Learn how to set up your Medical ID on smartphones now to ensure quick access when it matters most.

Installing Bed Exit Alarms

Installing bed exit alarms correctly can enhance safety, but ensuring optimal setup requires careful steps you won’t want to miss.

10 Smart Home Technologies That Will Revolutionize Senior Living!

Keep your loved ones safe and connected with innovative smart home technologies that can transform senior living—discover the top 10 game-changers inside!