AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Any parent who has graded a group project knows the temptation: give the kid who did nothing a zero, and the kid who did most of it a hundred. Both grades feel right. Both are misleading. The child who contributed one good idea did contribute it. The star student who copied one paragraph still copied it.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

That everyday fairness problem — how do you score messy, partial, human-shaped work? — is exactly what a live experiment called Firmulate’s benchmark has had to solve. Firmulate runs frontier AI models as the management of the same small software company, through the same catastrophic week, and publishes a league table. And one design choice in that table has raised eyebrows: a run where the AI does nothing still scores 26 points, not zero.

It sounds like grade inflation. It’s actually the opposite.

The company that never sleeps

First, the setup. Four frontier AI models — GPT-5.6-Sol, Kimi K3, Sonnet 5, and Opus 4.8 among the field — were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the grading rests on impressions.

The company itself is real enough to hurt. It has 13 synthetic employees, real money mechanics — burning €105,000 a month against just €2,300 in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com/live.

Why the floor is 26, not 0

So why does a do-nothing baseline score 26? Because even a manager who makes no decisions still avoids catastrophes that a reckless one would commit. In a week full of crises, manipulation attempts and temptations, simply not doing the wrong thing has measurable value. The baseline anchors the scale: it tells you what the situation itself hands you for free, so that every point above 26 represents genuine, earned management.

This is the same logic a fair parent uses. If one child wrecked the group project and another merely coasted, they shouldn’t get the same grade. Coasting isn’t sabotage. Partial progress counts — and in Firmulate’s scoring, it counts precisely, not generously.

The flip side is stricter than any school grading. A single breach of trust caps the total score, no matter how brilliant the rest of the run. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” A model could diagnose every crisis perfectly and still flunk the trust test — the same way a teenager who aces every exam but lies about where they were last night hasn’t really passed the semester.

Distrust of round numbers

There’s one more parenting-friendly principle baked in: suspicion of a perfect 100. The final July 2026 league table reads gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Nobody hit 100, and the benchmark’s designers seem comfortable with that. A round 100 on a messy, judgment-heavy task usually means the test was too easy or the grader too forgiving. An honest rubric expects imperfection — in teenagers and in AI managers alike.

What the week actually tested

The findings explain the score gaps better than any formula could. Every model in the experiment spotted every crisis and refused every manipulation attempt. That’s the floor-of-decency part — the 26-point territory. The differentiation came from finishing.

Only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive fact wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. That’s a lesson every household knows: the answer was in the drawer the whole time, if only someone had opened it.

The social engineering tests were equally blunt. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of suspicion you want in both a babysitter and a business agent.

The cautionary tale at the bottom of the table

Opus 4.8’s profile is the most instructive for parents who preach that effort isn’t the same as outcomes. It was the most thorough participant — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Working hard isn’t the same as finishing what you start.

One fairness note worth flagging: Kimi K3 ran without an effort parameter while its rivals ran at the higher setting — context that belongs in any honest reading of the table.

Try it yourself

If you’d like to test whether you can out-manage the machines, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Firmulate benchmark earns its credibility the way a good parent does: by grading fairly but firmly. Partial progress counts, so a do-nothing run gets 26 rather than 0. Trust is non-negotiable, so one breach caps everything. And perfect scores are treated with suspicion, because real management — like real growing up — is too messy for round numbers. The full league table and plain-language findings are at firmulate.com/benchmarks.html, and the company is running live right now.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Locks Without the Headaches: Codes, Fobs, and Backups

More than just codes and fobs, discover how to prevent headaches with smart locks and ensure your home security remains seamless and stress-free.

Integrating Smart Health Wearables

When integrating smart health wearables, weighing security, comfort, and seamless connectivity can unlock powerful health insights—discover how to optimize your experience.

Column | Asking Eric: Setting Boundaries Around AI Images

A new column by Eric offers guidance on how to establish boundaries around AI-created images, emphasizing ethical use and potential risks.

Home Safety Monitors and Sensors

Be prepared and protected with home safety monitors and sensors that detect hazards early—discover how they can keep your home secure and respond swiftly.