A solar company does not learn whether its systems can handle a rough week from a polished demo. A delayed installation, an unhappy customer or a tempting shortcut can reveal more. Firmulate has put several frontier AI models through that kind of test: running the same small software company amid crises, customer pressure and opportunities to cut corners.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
The experiment is part of Firmulate, a live, watchable simulation designed to measure how AI models manage a business. Its synthetic company has 13 employees and real money mechanics: monthly costs of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. The company runs every business day, and its decisions are versioned and auditable.
In the final July 2026 Crucible league, Moonshot’s Kimi K3 placed second with 93 points, behind gpt-5.6-sol at 95. It beat Sonnet 5, which scored 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total. Firmulate’s principle is blunt: “no amount of good work outweighs a breach of trust.”
The headline result is not simply that K3 ranked highly. All the models spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing the right move and completing it is the kind of weakness a fluent chat exchange can hide.
AI customer support software for solar companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The detail that changed the deal
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that buried security needle, closed the deal and saved the churning customer. It resisted all three baits and made one deviation, the cleanest discipline in the field.
That result has a practical parallel for solar and backup-power businesses. A model helping with customer support or operations needs to do more than respond convincingly to the latest message. It may need to find relevant context in company records, protect a customer relationship and follow through on a decision. Firmulate’s test asks whether an AI workforce can do those things under pressure.
AI business simulation tools for solar operations
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Restraint is part of the job
The models faced fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All five refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation was common across the field; finishing the work was not. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four. More analysis alone did not guarantee a complete result.
There is a fairness caveat to the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are useful evidence from one shared exercise, with that difference in mind. They do not establish how every model will perform in every company.
Firmulate publishes the benchmark and plain-language findings, and the live company can be watched. A quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each one.

AI document analysis software for customer records
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before you hire
For a business considering AI in its CRM, support queue or forecasts, the league makes one case clearly: model choice is a decision to test against your own work. A top score in a shared simulation is informative, but a company’s customers, records and risks are its own. Firmulate says enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making tools for renewable energy businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
