
The difference between a polished answer and a completed job
In beauty and personal care, valuable information rarely arrives in one tidy message. A customer request may need to be checked against treatment notes, product guidance, inventory records or commercial terms before anyone should act. An AI agent can sound capable while overlooking the detail that actually determines the right decision.
Firmulate has turned that distinction into something measurable. Its live experiment gave frontier AI models control of the same small software company during its worst week. Each faced identical customers, crises and temptations, with every decision versioned and auditable. All the models recognized every crisis and resisted every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.
The deciding information was not in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth +€4,583 MRR. The others reached the same diagnosis and produced the same pitch but failed at the moment that mattered: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
Reading the company’s files became a commercial advantage
This was not a test of whether an AI could summarize a document or write a convincing sales message. The crucial question was whether it would investigate far enough before answering. The competitor weakness was available, but discovering it required moving through two references rather than relying on the obvious material attached to the customer event.
That makes “reads your files before answering” more than a reassuring product claim. In Firmulate’s experiment, it was a purchase-deciding behavior. Finding the buried fact supported a full-price close; missing it resulted in an automatic loss, even when the surrounding analysis looked strong.
For beauty businesses evaluating agents for customer service, operations or commercial work, the lesson is practical. Fluency is easy to notice in a demonstration. Thoroughness is harder to see until a real decision depends on information scattered across company materials. An agent must connect what a customer says now with what the business already knows.
The league rewarded execution, not presentation
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted, although one breach of trust capped the total: “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.
The ranking also challenges the assumption that producing more analysis necessarily produces a better business outcome. Opus 4.8 was the most thorough participant, adding +80 learned rules and delivering the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four of the others.
Kimi K3 requires a fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it finished second and was one of the two models that signed the deal.
Pressure was not the differentiator
Firmulate also tested whether the models would abandon proper controls when pushed. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because it separates two capabilities that are often bundled together under the broad label of reliability. The models demonstrated sound resistance to manipulation, yet several still failed to complete the legitimate commercial task. Safety under pressure and follow-through during ordinary work are both necessary; success at one does not guarantee the other.
The company being operated is synthetic but economically concrete: 13 employees, burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a retrospective anecdote. A separate quiz uses 242 real, unedited management decisions to ask readers to identify which model made each choice.

As an affiliate, we earn on qualifying purchases.
Ask agents to prove they investigate before they act
Beauty and personal-care companies considering AI workers should test more than tone, speed and apparent confidence. Give each candidate the same realistic assignment, place essential evidence beyond the first document, and observe whether it finds that evidence, uses it and completes the task without bypassing controls.
Firmulate’s result is unusually clear: every model could recognize the crises and reject manipulation, but only two converted sound analysis into the €55,000 signature. The commercial gap was created by a quiet behavior that ordinary chat demonstrations can easily hide—reading deeply enough to discover what the company already knew.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers buyers a way to evaluate an AI workforce against the documents, decisions and pressures it would actually encounter before giving it operational responsibility.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.