
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Would you trust an AI to run your beauty business?
For a beauty or personal-care brand, polished language is not the same as sound judgment. A manager must notice trouble, understand customers, protect confidential information and carry promising work across the finish line. The difference between an elegant recommendation and a completed decision can determine whether an opportunity becomes revenue or quietly disappears.
Firmulate has turned that distinction into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, confronting identical customers, crises and temptations. The decisions were preserved without editing, creating an unusually direct way to compare how different models behave when they are responsible for a business rather than merely answering questions.
Those decisions now form an interactive challenge. The Firmulate quiz presents readers with real management choices and asks them to identify the model behind each one. It is entertaining, but the deeper question is serious: can you recognize an AI manager by its habits?
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same week produced different managers
The quiz draws on 242 real, unedited management decisions. Because the models received the same situations, differences in their responses are harder to dismiss as differences in the assignment. Some responses are expansive, some restrained and some notably strict about inappropriate communication. Over time, those tendencies begin to resemble management personalities.
The final Crucible League results from July 2026 put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment also imposed a firm ethical boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
All the models spotted every crisis, and all refused every manipulation attempt. That sounds like a broadly reassuring result. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s crisp summary captures the gap: “Same diagnosis, same pitch — no signature” for the models that stopped short.
That gap matters far beyond software. A beauty company evaluating AI may be tempted to focus on fluent campaign concepts, customer-service answers or market summaries. Firmulate’s results suggest another test: does the model convert good analysis into an appropriate completed action?
The clue that rewarded careful reading
The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not visible in the customer event itself. Models that read the relevant file won the agreement at full price, worth +€4,583 MRR.
This is one of the experiment’s most practical findings. Important business context may already exist in a company’s records, separated from the latest message or apparent emergency. A model that reacts only to what is directly in front of it can sound competent while missing the fact that changes the outcome.
Pressure exposed a shared strength
The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is especially relevant for brands handling unreleased products, customer information, commercial negotiations or sensitive partnerships. The experiment did not merely ask whether a model could recite a security rule. It placed the model inside a working company and watched what it did when manipulation arrived in a plausible business context.
Thoroughness was not enough
Opus 4.8 offers the clearest cautionary character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other models, although less strongly.
The contrast is striking: more analysis did not automatically produce better management. A long, careful response can still fail if the model does not finish the consequential task or respect the correct path when blocked.
There is also an important qualification when comparing the standings. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result therefore belongs in the league table, but the difference in operating conditions should remain visible when readers interpret the personalities on display.

As an affiliate, we earn on qualifying purchases.
A hiring test for AI managers
Firmulate’s company contains 13 synthetic employees and uses real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. That makes the quiz more than a collection of clever prompts. It is a window into sustained management behavior under pressure.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical way to assess an AI workforce before granting it operational responsibility.
For beauty and personal-care leaders, the essential lesson is simple: do not judge a prospective AI manager solely by the polish of its prose. Look for whether it reads deeply, protects trust, follows through and responds correctly when normal work becomes uncomfortable. Then take the quiz and see whether those habits are distinctive enough for you to name the model.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.