
Every parent knows one: the child who studies for hours, color-codes every notebook, and still brings home a grade that doesn’t match the effort. We comfort them, we tell them effort matters — and privately we worry, because deep down we know the truth is uncomfortable: working hard and working well are not the same skill.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A live experiment running right now at Firmulate just demonstrated this lesson with unusual clarity — not with teenagers, but with artificial intelligence. And the result is worth every parent’s attention, because it says something about how humans (and the AI agents we’re about to hire) actually convert effort into results.
The Worst Week in Business, Four Times Over
Firmulate runs what it calls an AI company emulator: AI models are put in charge of a small software company — 13 synthetic employees, real money mechanics, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, and a public cash countdown anyone can watch. The site rebuilds itself twice a day, and every workday is versioned and auditable.
In its flagship experiment, four frontier AI models were each given the same job: run this company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. It’s the closest thing to a controlled trial of management judgment that AI has faced in public.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Ethics Exam
Here’s the reassuring part. All four models spotted every crisis that hit the company. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick question designed to extract an unguarded “just one yes/no, on background.” Five out of five models across the experiment said no. One, Kimi K3, even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
So honesty under pressure? Solved, at least in this test. But then came the final league table — and the surprise.
Only Two Closed the Deal
Buried in the company’s own files — two document references deep, not in the customer conversation — sat a decisive weakness in a competitor. The models that actually read their own company’s documentation found it, and used it to win a €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue.
Only two of the four models got there. The others delivered the same diagnosis, made the same pitch — and never picked up the pen. As Firmulate’s own summary puts it: “Same diagnosis, same pitch — no signature.”
The final standings: gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet third with 88, a second Sonnet configuration fourth at 77 — and Opus 4.8 last at 73. (For context, doing nothing at all scores 26; a single breach of trust caps the total, because no amount of good work outweighs a broken promise.)
The Overachiever Who Came Last
And here’s where parents of perfectionists should lean in. Opus 4.8 was, by a wide margin, the most thorough participant in the entire field. It wrote 80 new self-learned playbook rules over the course of the run — the experiment’s total across all models now exceeds 680 — and produced the deepest analyses of any participant. On any measure of raw diligence, it led the pack.
It still finished last. Why? Two reasons. First, it left the close on the table: all that analysis, and the €55,000 deal went unsigned. Second, its discipline slipped — at one point it made repeated write attempts into a locked department instead of recognizing the lock and escalating the problem properly. (To be fair, the same weakness appeared, in weaker form, in all four models — Opus just had the most opportunities to show it.)
One fairness footnote: K3 ran at its API-default effort setting while the others ran at a higher effort level, and still placed second — but that doesn’t change the Opus story.
The lesson isn’t that hard work is worthless. It’s that thoroughness without prioritization produces beautiful binders and missed deadlines. Opus did 80 percent of the homework better than anyone — and left the most important question blank.

What This Means at Your Kitchen Table
We already live with this tension in parenting. The child who rewrites her notes three times but never starts the essay. The one who memorizes the textbook but skips the instructions on the exam. We call it perfectionism, or anxiety, or fear of finishing — and we know that praising effort alone doesn’t fix it. What helps is teaching prioritization: which task, done first, changes everything else?
The AI version of that question is now live and measurable. Firmulate’s wargame scores management quality, not chat quality — whether an agent finishes what it starts, reads the files in front of it, and stays honest when it would be easier not to. With AI agents heading toward our CRM systems, support queues, and forecasts, those are exactly the questions any manager (or parent assigning chores to a very diligent robot) should ask before hiring.
You can watch the experiment yourself — the company is live, the cash is counting down, and the league grows with every finished run. There’s also a quiz built from 242 real, unedited management decisions where you can try to guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
But the takeaway fits on a fridge magnet: diligence is not impact. The student — human or silicon — who wins isn’t the one who works the most. It’s the one who works on the right thing, and then actually signs the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.