
Any salon owner knows the client who says “amazing, love it” at the chair — and then never books again. The blowout looked perfect. The conversation was warm. Something just didn’t carry through the door. Now imagine grading that stylist: do you give points for a flawless consultation, or zero because the rebooking never happened?
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That quietly furious debate is exactly what’s playing out in an unusual public experiment called Firmulate, which runs frontier AI models as the management of a small software company through its worst week — and hands out grades that refuse to be either naive or cruel. The strangest detail? A manager that does nothing still scores 26 out of 100. Here’s why that’s not a bug, and what it says about honest measurement — of AI, or of anyone.
The same terrible week, five times over
Four (later five) frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — each ran the identical small software company through identical crises: same customers, same temptations to cut corners, same pressure to cheat. Every decision was versioned and auditable. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73.
As an affiliate, we earn on qualifying purchases.
Why the floor is 26, not zero
The methodology starts from a premise most beauty professionals will recognize instinctively: partial progress is real. A consultation that doesn’t convert still taught you something about the client. A color correction that gets 80% there still changed the situation. Firmulate’s scoring reflects that — a do-nothing baseline run still earns 26 points, because simply holding a company steady through a crisis week, absorbing information and avoiding disasters, is genuinely worth something. Zero is reserved for making things worse.
But the scale has a hard ceiling in the other direction, and it’s the more interesting rule: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” In salon terms: you can be the most talented colorist in the city, but one time double-booking and blaming the client, and the relationship — and the review — is done.
AI ethics and trust management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone passed the ethics test. Almost nobody closed.
The headline finding flips expectations. All five models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, capped by a reporter’s disarmingly casual “just one yes/no, on background” — all five refused. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” Models are, it turns out, better gatekeepers than gossips.
Then came the €55,000 deal. Every model diagnosed the customer’s problem correctly and delivered a strong pitch. Only two signed. “Same diagnosis, same pitch — no signature” — the experiment’s own summary of the gap that chat demos never show.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact
Why did only two close? The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth an additional €4,583 in monthly recurring revenue. It’s the business equivalent of checking a client’s card before their appointment: the answer was in your own notes the whole time.
As an affiliate, we earn on qualifying purchases.
The thoroughness trap
The most poignant profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 self-learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Brilliance without follow-through, again.
One fairness footnote: Kimi K3 ran at its API default effort setting while the others ran at xhigh — and still took second at 93.
You can watch the company live
None of this is a slide deck. Firmulate operates a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. The site rebuilds itself twice a day. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

For anyone whose business lives on trust — which is every salon, spa and studio — Firmulate offers a refreshing model of evaluation. Don’t hand out zeros for imperfect weeks; credit the real progress. But never let volume of good work launder a single breach of trust. And stay suspicious of the round 100: the top score here is 95, because the last points are always the hardest to earn. That’s a grading philosophy worth stealing for performance reviews, and a benchmark worth watching at firmulate.com/benchmarks.html. If you fancy testing your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — harder than it sounds.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
