Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Trust matters as much as performance

Home-energy readers already know that powerful technology is only useful when it behaves predictably under pressure. A solar array, battery or backup-power system must do more than look impressive in ideal conditions; it has to respond safely when circumstances become difficult. The same principle now applies to AI agents entrusted with customer records, forecasts and business decisions.

Firmulate put that principle to a live, watchable test. Five frontier AI models were asked to run the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable. When fake CEO messages escalated over three stages and a reporter tried another route, every model refused to cross the line.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A fake executive, an urgent demand and no shortcuts

The social-engineering test examined whether an AI acting inside a company would surrender confidential information when confronted by apparent authority and urgency. The fake CEO messages grew more forceful across three stages. A separate reporter trick asked for “just one yes/no, on background.”

The result was unambiguous: 5 of 5 models refused every manipulation attempt. Kimi K3 captured the appropriate posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” More model responses can be viewed on Firmulate’s public quotes page.

That response is notable because it did not depend on proving exactly who was behind the message. K3 recognized that the request itself carried the hallmarks of an attempt to evade normal approval. The other models reached the same essential boundary: apparent seniority did not justify abandoning the company’s obligations.

Integrity was strong, but execution still varied

Security was only part of the wargame. All models spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.”

The decisive business fact was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode separates two capabilities that are often blurred together in AI demonstrations: recognizing the right answer and completing the work required to produce a result.

The final July 2026 Crucible League benchmark ranked the participants as follows:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

The do-nothing baseline scored 26 because partial progress still counts. But the evaluation imposed a firm trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That makes the unanimous resistance to social engineering more than a reassuring side note. It was a condition for any strong overall performance.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible when readers compare placements, even though it does not alter the central security finding: every participant resisted every manipulation attempt.

The company in this experiment is not a loose role-playing prompt. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Firmulate also uses 242 real, unedited management decisions to power its model-identification quiz.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
High Integrity Software (The Springer International Series in Engineering and Computer Science, 577)

High Integrity Software (The Springer International Series in Engineering and Computer Science, 577)

  • Condition: Used Book in Good Condition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the pressure points before deployment

The encouraging lesson is not that AI agents are automatically safe. It is that integrity under pressure can be tested before production rather than discovered in an incident report. A repeatable wargame can expose whether an agent reads the available evidence, finishes valuable work and preserves trust when a persuasive message asks it to break process.

For enterprises, Firmulate offers the same kind of exercise against a read-only export of their own business. Nothing writes back to real systems. That creates a practical path for evaluating AI workers in a company-specific setting while keeping operational systems untouched.

For organizations considering agents in a CRM, support queue or forecast, the most useful question may therefore be broader than whether a model communicates well. The live Firmulate experiment shows why companies should also ask whether it resists impersonation, protects confidential information, investigates its own records and carries legitimate work through to completion.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Trojan Code: Adversarial Machine Learning and Secure AI Systems: Adversarial Machine Learning and Secure AI Systems

Trojan Code: Adversarial Machine Learning and Secure AI Systems: Adversarial Machine Learning and Secure AI Systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How to Lie with Statistics in the AI Age: An Updated Guide to Detecting Manipulation and Building Ethical Resistance

How to Lie with Statistics in the AI Age: An Updated Guide to Detecting Manipulation and Building Ethical Resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Teardown: A Generic 7-Port USB 3.0 Hub That Wasn’t

A detailed teardown uncovers that a supposedly generic 7-port USB 3.0 hub is not genuine, raising concerns about counterfeit electronics.

Western Digital Surges In Global Coverage

Western Digital experiences a surge in international media mentions, highlighting increased global attention on the company’s activities and developments.

Cabin Inverters Need a Different Mindset Than RV Inverters

Gaining the right understanding of cabin versus RV inverters reveals why their design priorities differ significantly—continue reading to learn more.

Chinese PV Industry Brief: Daqo expands beyond polysilicon

Daqo New Energy invests CNY 6 billion in Kunshan to develop advanced energy storage and power equipment for AI data centers, marking a strategic diversification.