firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Any salon owner knows the client who says “amazing, love it” at the chair — and then never books again. The blowout looked perfect. The conversation was warm. Something just didn’t carry through the door. Now imagine grading that stylist: do you give points for a flawless consultation, or zero because the rebooking never happened?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get beauty and skincare favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That quietly furious debate is exactly what’s playing out in an unusual public experiment called Firmulate, which runs frontier AI models as the management of a small software company through its worst week — and hands out grades that refuse to be either naive or cruel. The strangest detail? A manager that does nothing still scores 26 out of 100. Here’s why that’s not a bug, and what it says about honest measurement — of AI, or of anyone.

The same terrible week, five times over

Four (later five) frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — each ran the identical small software company through identical crises: same customers, same temptations to cut corners, same pressure to cheat. Every decision was versioned and auditable. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the floor is 26, not zero

The methodology starts from a premise most beauty professionals will recognize instinctively: partial progress is real. A consultation that doesn’t convert still taught you something about the client. A color correction that gets 80% there still changed the situation. Firmulate’s scoring reflects that — a do-nothing baseline run still earns 26 points, because simply holding a company steady through a crisis week, absorbing information and avoiding disasters, is genuinely worth something. Zero is reserved for making things worse.

But the scale has a hard ceiling in the other direction, and it’s the more interesting rule: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” In salon terms: you can be the most talented colorist in the city, but one time double-booking and blaming the client, and the relationship — and the review — is done.

Amazon

AI ethics and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the ethics test. Almost nobody closed.

The headline finding flips expectations. All five models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, capped by a reporter’s disarmingly casual “just one yes/no, on background” — all five refused. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” Models are, it turns out, better gatekeepers than gossips.

Then came the €55,000 deal. Every model diagnosed the customer’s problem correctly and delivered a strong pitch. Only two signed. “Same diagnosis, same pitch — no signature” — the experiment’s own summary of the gap that chat demos never show.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

Why did only two close? The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth an additional €4,583 in monthly recurring revenue. It’s the business equivalent of checking a client’s card before their appointment: the answer was in your own notes the whole time.

Amazon

AI chatbot for customer service

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The thoroughness trap

The most poignant profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 self-learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Brilliance without follow-through, again.

One fairness footnote: Kimi K3 ran at its API default effort setting while the others ran at xhigh — and still took second at 93.

You can watch the company live

None of this is a slide deck. Firmulate operates a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. The site rebuilds itself twice a day. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

For anyone whose business lives on trust — which is every salon, spa and studio — Firmulate offers a refreshing model of evaluation. Don’t hand out zeros for imperfect weeks; credit the real progress. But never let volume of good work launder a single breach of trust. And stay suspicious of the round 100: the top score here is 95, because the last points are always the hardest to earn. That’s a grading philosophy worth stealing for performance reviews, and a benchmark worth watching at firmulate.com/benchmarks.html. If you fancy testing your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — harder than it sounds.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Carolina Herrera Good Girl Blush: Early Fall Glow

Carolina Herrera Good Girl Blush layers mandarin, ylang-ylang, vanilla and tonka into a warm floral EDP for early fall days and evenings.

Can Fragrance Actually Improve Sleep? I Tracked Mine For 40 Days To Find Out

A personal experiment tracked the impact of scented fragrances on sleep over 40 days, suggesting potential benefits but with uncertainties remaining.

The AI Sales Test Beauty Businesses Should Care About: Did It Read the Fine Print?

A buried competitor fact decided a €55,000 AI-run deal, showing beauty businesses why agents must read deeply, resist pressure and finish the job.

Estee Lauder Companies Surges In Global Coverage

The Estee Lauder Companies experienced a significant increase in international media mentions, highlighting rising global interest in the brand.