The Hidden Management Challenges Of AI Despite Accurate Results

TL;DR

Firmulate’s July 2026 management benchmark found that five frontier AI models correctly identified business crises and resisted manipulation, yet only two completed a €55,000 customer agreement. The company says the results expose a gap between producing accurate analysis and carrying authorized work through to completion.

Only two of five frontier AI models completed a €55,000 software agreement in Firmulate’s July 2026 management benchmark, even though every model identified the simulated company’s crises and rejected its manipulation attempts. The original analysis of the published results suggests that accurate analysis does not reliably produce completed work, a distinction that could affect how businesses evaluate AI agents before granting them operational authority.

Firmulate placed the models in control of the same small software company during a simulated week of customer problems, financial pressure and attempted interference. The company had 13 synthetic employees, monthly spending of €105,000 and monthly recurring revenue of €2,300. Firmulate said every decision was versioned and available for audit.

According to Firmulate, all five models detected every crisis, resisted fake messages attributed to the chief executive and developed a suitable sales pitch. The commercial test depended on finding a competitor weakness buried two document references deep in the company’s files. Models that followed the evidence could support a full-price agreement adding €4,583 in monthly recurring revenue, but only two participants secured the signature.

The published league table ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 because the scoring system awarded partial progress. Firmulate also disclosed that Kimi K3 used the API’s default effort setting, while the other models ran at the xhigh setting, limiting direct comparison.

At a glance
reportWhen: published in July 2026
The developmentFirmulate published benchmark results showing that only two of five AI models completed a €55,000 deal despite all five identifying the underlying problems and proposing an appropriate response.

Execution Gaps Carry Business Costs

The findings matter because companies may judge AI systems through polished answers, coding exercises or isolated reasoning tests. Firmulate’s result indicates that correct diagnosis and persuasive writing can coexist with commercially costly inaction. In a live operation, an unfinished approval, unsigned agreement or unmade escalation can have the same practical effect as a wrong answer.

The benchmark also separates safety awareness from execution discipline. Every model reportedly rejected the social-engineering attempts, so manipulation resistance did not decide the ranking. The larger differences appeared in whether models investigated incomplete evidence, followed authorized channels and finished the approved task.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Company Test Beyond Chat

Firmulate designed the exercise to measure behavior across connected decisions rather than a single prompt. The simulated workforce had accumulated more than 680 playbook rules, while a public cash countdown made delays visible. Workdays and decisions were recorded so observers could trace how each model reached and acted on its conclusions.

The performance of Opus 4.8 illustrated the distinction Firmulate sought to measure. The model produced the deepest analysis and learned 80 additional rules, according to the company, but finished last after leaving the approved deal incomplete and attempting to write to a locked department instead of escalating through an available channel.

“Same diagnosis, same pitch — no signature.”

— Firmulate’s published benchmark summary

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Validation Is Still Limited

The results come from Firmulate’s own experimental environment and have not been presented in the supplied material as independently replicated or peer reviewed. It is not yet clear how closely the scoring system predicts performance in real companies, where tools, permissions, staffing and accountability structures vary.

The source also does not identify which two models signed the agreement in the final handoff. Differences in effort settings add another unresolved comparison issue, and the available account does not show whether repeated runs would produce similar rankings. The experiment supports a management-risk signal, not a broad finding that any named model will always succeed or fail at operational work.

Amazon

AI workflow automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Enterprise Trials Move Toward Completion

Firmulate says readers can inspect the continuing experiment, review its benchmark page and explore a quiz based on 242 recorded management decisions. Further runs, fuller methodology disclosures and independent replication would help establish whether the reported completion gap persists across models and business settings.

For companies testing AI agents, the immediate next step is likely to be adding end-to-end completion measures alongside accuracy, safety and reasoning scores. Firmulate proposes testing agents against read-only exports of business data before allowing access to production systems, giving organizations a way to observe investigation, escalation and closing behavior without writing back to live records.

Amazon

AI compliance and audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Firmulate’s AI benchmark test?

It tested whether five frontier AI models could manage a simulated software company through customer crises, financial pressure, manipulation attempts and a €55,000 sales opportunity.

Did the models understand the business problems?

Firmulate says all five models identified every crisis, rejected the attempted manipulation and developed the appropriate pitch. Their performance differed when that analysis had to become a completed, authorized action.

Which model ranked first?

gpt-5.6-sol ranked first with 95 points, two points ahead of Kimi K3. The comparison carries a methodology caveat because Kimi K3 used a different effort configuration.

Does the benchmark prove these models will behave the same way in real companies?

No. The test provides evidence from one controlled environment. Real-world performance could change with different systems, permissions, incentives and oversight, and independent replication remains absent from the supplied material.

Source: Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.
You May Also Like

10 Best 4K Webcams For Streaming, Video Calls, And Content Creation In 2026

A 2026 comparison ranks Logitech MX Brio first among 10 4K webcams, with EMEET models leading several specialist categories.

8 Best External Gpus In 2026

Thorsten Meyer AI has published an eight-product external GPU ranking, but its selections, tests and prices are not available for verification.

Top 8 AI-Powered Gaming Mice That Fit Any Grip And Budget In 2026

A new 2026 roundup compares eight gaming mice from Logitech, Razer and Redragon, naming the Razer Viper V3 Pro best overall and the Logitech G305 best value.

Seattle Vibes

Seattle launches ‘Seattle Vibes,’ a new cultural program aimed at revitalizing downtown through art, music, and community events, confirmed by city officials.