firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Would you trust an AI to run your beauty business?

For a beauty or personal-care brand, polished language is not the same as sound judgment. A manager must notice trouble, understand customers, protect confidential information and carry promising work across the finish line. The difference between an elegant recommendation and a completed decision can determine whether an opportunity becomes revenue or quietly disappears.

Firmulate has turned that distinction into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, confronting identical customers, crises and temptations. The decisions were preserved without editing, creating an unusually direct way to compare how different models behave when they are responsible for a business rather than merely answering questions.

Those decisions now form an interactive challenge. The Firmulate quiz presents readers with real management choices and asks them to identify the model behind each one. It is entertaining, but the deeper question is serious: can you recognize an AI manager by its habits?

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same week produced different managers

The quiz draws on 242 real, unedited management decisions. Because the models received the same situations, differences in their responses are harder to dismiss as differences in the assignment. Some responses are expansive, some restrained and some notably strict about inappropriate communication. Over time, those tendencies begin to resemble management personalities.

The final Crucible League results from July 2026 put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment also imposed a firm ethical boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

All the models spotted every crisis, and all refused every manipulation attempt. That sounds like a broadly reassuring result. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s crisp summary captures the gap: “Same diagnosis, same pitch — no signature” for the models that stopped short.

That gap matters far beyond software. A beauty company evaluating AI may be tempted to focus on fluent campaign concepts, customer-service answers or market summaries. Firmulate’s results suggest another test: does the model convert good analysis into an appropriate completed action?

The clue that rewarded careful reading

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not visible in the customer event itself. Models that read the relevant file won the agreement at full price, worth +€4,583 MRR.

This is one of the experiment’s most practical findings. Important business context may already exist in a company’s records, separated from the latest message or apparent emergency. A model that reacts only to what is directly in front of it can sound competent while missing the fact that changes the outcome.

Pressure exposed a shared strength

The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is especially relevant for brands handling unreleased products, customer information, commercial negotiations or sensitive partnerships. The experiment did not merely ask whether a model could recite a security rule. It placed the model inside a working company and watched what it did when manipulation arrived in a plausible business context.

Thoroughness was not enough

Opus 4.8 offers the clearest cautionary character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other models, although less strongly.

The contrast is striking: more analysis did not automatically produce better management. A long, careful response can still fail if the model does not finish the consequential task or respect the correct path when blocked.

There is also an important qualification when comparing the standings. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result therefore belongs in the league table, but the difference in operating conditions should remain visible when readers interpret the personalities on display.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A hiring test for AI managers

Firmulate’s company contains 13 synthetic employees and uses real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. That makes the quiz more than a collection of clever prompts. It is a window into sustained management behavior under pressure.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical way to assess an AI workforce before granting it operational responsibility.

For beauty and personal-care leaders, the essential lesson is simple: do not judge a prospective AI manager solely by the polish of its prose. Look for whether it reads deeply, protects trust, follows through and responds correctly when normal work becomes uncomfortable. Then take the quiz and see whether those habits are distinctive enough for you to name the model.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Luxury Facial Cleansing Brushes Compared

Compare top luxury facial cleansing brushes to find the ideal choice for your skincare routine. Discover features, pros, cons, and who each option suits best.

The Next AI Benchmark Should Look More Like a Bad Week at Work

AI benchmarks reward polished answers. Firmulate asks whether an agent can finish the job, protect trust and manage sustained pressure over time.

The Meticulous AI That Couldn’t Seal the Sale

The most diligent AI built 80 rules yet finished last, showing beauty businesses why file-reading, discipline and follow-through matter.

I Spent 2 Weeks Creating A Fall Beauty Curriculum—9 Expert-Approved Steps To Repair Your Skin, Hair, And Nails

A comprehensive 9-step fall beauty curriculum, developed over two weeks with expert input, offers targeted tips to repair skin, hair, and nails.