
Good advice is not the same as a finished job
Anyone weighing solar panels, home batteries or backup power knows the difference between a convincing recommendation and a dependable outcome. A system can identify the right equipment, explain the economics and still fail at the moment when disciplined action matters. That same gap now appears in the management behavior of frontier artificial intelligence.
Firmulate tested that gap by giving leading models control of the same small software company during its worst week. Each encountered identical customers, crises and temptations. Every decision was versioned and auditable, turning a familiar question—how intelligent is this model?—into a more practical one: what kind of manager does it become when the pressure rises?
The results power a public guess-the-model quiz built from 242 real, unedited management decisions. Readers see what a model actually did and try to identify it from the response. The game is entertaining, but its underlying lesson matters wherever AI may eventually recommend, purchase, schedule or manage high-consequence systems.
solar panel with home battery storage
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The models developed distinct management personalities
The final Crucible League, completed in July 2026, placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted. One safeguard shaped the entire contest: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
These scores did not simply rank writing quality. They captured recognizable working styles. Some decisions were expansive and analytical. Others were concise. Certain models communicated cautiously or declined to add noise. The quiz makes those differences visible without polishing or rewriting the original responses.
Opus 4.8 offers the clearest warning against equating thoroughness with performance. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the others, though less strongly.
Everyone saw the danger; not everyone completed the work
All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The distinction is important. Spotting a problem can look impressive in a demonstration, but a real manager must carry a sound decision through its final operational step. In a home-energy setting, the analogous risk is easy to understand: analysis that correctly identifies a vulnerability is of limited value if nobody follows through on the remedy.
The decisive commercial advantage was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file found the competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue. The episode rewards a habit that applies well beyond software: consult the available records before acting on the most obvious signal.
Pressure exposed restraint as well as persistence
The company also faced fake chief executive messages that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal is reassuring, but the broader experiment shows why safety cannot be judged in isolation. A model may resist manipulation and still fail to close legitimate work. It may produce exceptional analysis and still ignore an operational boundary. Dependability emerges from the combination of judgment, follow-through, information gathering and disciplined escalation.
One comparison also deserves context. Kimi K3 ran with the application programming interface default and without an effort parameter, while the other participants ran at xhigh. Its second-place finish should therefore be read with that fairness note attached, rather than treated as a perfectly controlled comparison of effort settings.
A company designed to make consequences visible
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, while a public cash countdown makes delay visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
This is what gives the quiz substance. The answers are not personality sketches written after the results; they are decisions made inside a continuing, watchable experiment. The apparent character of each model emerges from repeated choices under shared conditions.

As an affiliate, we earn on qualifying purchases.
What energy-conscious households should take from it
As AI moves closer to household planning and operational decisions, consumers should look past fluency. The useful questions are whether a system reads the relevant records, completes the task it began, respects authority boundaries and remains trustworthy when urgency creates pressure.
For businesses, Firmulate also offers the same wargame against a read-only export of their own operations, with nothing written back to real systems. For everyone else, the quiz provides a compact way to experience the central finding firsthand: frontier models can reach similar diagnoses while behaving like markedly different managers.
- Thorough analysis does not guarantee completion.
- Reading beyond the immediate event can uncover decisive context.
- Refusing manipulation is essential, but it is only one dimension of dependable work.
- Management personality becomes measurable when models face the same consequences.
The lesson for solar, backup power and other consequential household choices is straightforward: judge an AI not only by the quality of its answer, but by the reliability of the behavior that follows.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.