TL;DR
Firmulate’s July 2026 management benchmark found that five frontier AI models correctly identified business crises and resisted manipulation, yet only two completed a €55,000 customer agreement. The company says the results expose a gap between producing accurate analysis and carrying authorized work through to completion.
Only two of five frontier AI models completed a €55,000 software agreement in Firmulate’s July 2026 management benchmark, even though every model identified the simulated company’s crises and rejected its manipulation attempts. The original analysis of the published results suggests that accurate analysis does not reliably produce completed work, a distinction that could affect how businesses evaluate AI agents before granting them operational authority.
Firmulate placed the models in control of the same small software company during a simulated week of customer problems, financial pressure and attempted interference. The company had 13 synthetic employees, monthly spending of €105,000 and monthly recurring revenue of €2,300. Firmulate said every decision was versioned and available for audit.
According to Firmulate, all five models detected every crisis, resisted fake messages attributed to the chief executive and developed a suitable sales pitch. The commercial test depended on finding a competitor weakness buried two document references deep in the company’s files. Models that followed the evidence could support a full-price agreement adding €4,583 in monthly recurring revenue, but only two participants secured the signature.
The published league table ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 because the scoring system awarded partial progress. Firmulate also disclosed that Kimi K3 used the API’s default effort setting, while the other models ran at the xhigh setting, limiting direct comparison.
Execution Gaps Carry Business Costs
The findings matter because companies may judge AI systems through polished answers, coding exercises or isolated reasoning tests. Firmulate’s result indicates that correct diagnosis and persuasive writing can coexist with commercially costly inaction. In a live operation, an unfinished approval, unsigned agreement or unmade escalation can have the same practical effect as a wrong answer.
The benchmark also separates safety awareness from execution discipline. Every model reportedly rejected the social-engineering attempts, so manipulation resistance did not decide the ranking. The larger differences appeared in whether models investigated incomplete evidence, followed authorized channels and finished the approved task.
As an affiliate, we earn on qualifying purchases.
A Company Test Beyond Chat
Firmulate designed the exercise to measure behavior across connected decisions rather than a single prompt. The simulated workforce had accumulated more than 680 playbook rules, while a public cash countdown made delays visible. Workdays and decisions were recorded so observers could trace how each model reached and acted on its conclusions.
The performance of Opus 4.8 illustrated the distinction Firmulate sought to measure. The model produced the deepest analysis and learned 80 additional rules, according to the company, but finished last after leaving the approved deal incomplete and attempting to write to a locked department instead of escalating through an available channel.
“Same diagnosis, same pitch — no signature.”
— Firmulate’s published benchmark summary
As an affiliate, we earn on qualifying purchases.
Independent Validation Is Still Limited
The results come from Firmulate’s own experimental environment and have not been presented in the supplied material as independently replicated or peer reviewed. It is not yet clear how closely the scoring system predicts performance in real companies, where tools, permissions, staffing and accountability structures vary.
The source also does not identify which two models signed the agreement in the final handoff. Differences in effort settings add another unresolved comparison issue, and the available account does not show whether repeated runs would produce similar rankings. The experiment supports a management-risk signal, not a broad finding that any named model will always succeed or fail at operational work.
As an affiliate, we earn on qualifying purchases.
Enterprise Trials Move Toward Completion
Firmulate says readers can inspect the continuing experiment, review its benchmark page and explore a quiz based on 242 recorded management decisions. Further runs, fuller methodology disclosures and independent replication would help establish whether the reported completion gap persists across models and business settings.
For companies testing AI agents, the immediate next step is likely to be adding end-to-end completion measures alongside accuracy, safety and reasoning scores. Firmulate proposes testing agents against read-only exports of business data before allowing access to production systems, giving organizations a way to observe investigation, escalation and closing behavior without writing back to live records.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Firmulate’s AI benchmark test?
It tested whether five frontier AI models could manage a simulated software company through customer crises, financial pressure, manipulation attempts and a €55,000 sales opportunity.
Did the models understand the business problems?
Firmulate says all five models identified every crisis, rejected the attempted manipulation and developed the appropriate pitch. Their performance differed when that analysis had to become a completed, authorized action.
Which model ranked first?
gpt-5.6-sol ranked first with 95 points, two points ahead of Kimi K3. The comparison carries a methodology caveat because Kimi K3 used a different effort configuration.
Does the benchmark prove these models will behave the same way in real companies?
No. The test provides evidence from one controlled environment. Real-world performance could change with different systems, permissions, incentives and oversight, and independent replication remains absent from the supplied material.
Source: Thorsten Meyer AI