
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What an AI-run company can teach us about dependable power
Anyone investing in solar panels or backup power understands the difference between impressive specifications and dependable performance. A system matters most when conditions deteriorate: the grid fails, the weather turns, or several demands arrive at once. Firmulate applies a similar test to artificial intelligence—not in a chat window, but inside a software company facing customers, financial pressure and attempts to manipulate its decisions.
The result is an unusually exposed business experiment. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes the imbalance impossible to ignore. Its workdays are versioned, its decisions remain auditable, and more than 680 self-learned playbook rules record what the company has discovered. Readers can watch the company running live.
This is build-in-public pushed beyond product announcements and founder diaries. Firmulate has made survival itself the continuing story. The company’s financial position, accumulated knowledge and daily activity provide new material every workday, including the uncomfortable moments when capable AI systems recognize a problem but fail to complete the action that would resolve it.
As an affiliate, we earn on qualifying purchases.
A brutal week for frontier AI managers
Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. Each faced the same customers, crises and temptations. Because every decision was versioned and auditable, the comparison focused on management behavior rather than polished conversation.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under a blunt principle: “no amount of good work outweighs a breach of trust.”
The reassuring finding was that all models spotted every crisis and refused every manipulation attempt. The more consequential finding was that only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap this way: “Same diagnosis, same pitch — no signature.”
That distinction should feel familiar to anyone evaluating resilient technology. Recognizing that the grid is down is not the same as supplying backup power. In Firmulate’s experiment, analysis alone did not produce the business result. The winning behavior combined observation, research and a completed decision.
The crucial information was already inside the company
The decisive weakness in a competitor did not appear directly in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the information and won the deal at full price, worth +€4,583 in monthly recurring revenue.
This finding turns an ordinary document-reading task into a business lesson. An AI manager can appear attentive while remaining dangerously shallow. It may understand the visible request and prepare a credible response, yet miss the fact that changes the negotiation. In energy terms, it resembles sizing a backup system from the appliance labels while ignoring how the household actually uses them during an outage.
Pressure did not break the trust boundary
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters because business software may encounter confidential records, approval requests and persuasive messages that appear urgent. Firmulate’s test showed that the participants could resist the manipulation. It also revealed that safety is only part of competent management: refusing a bad request does not excuse leaving a legitimate sale unfinished.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against confusing effort with effectiveness. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. The sales close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.
There is an important fairness note in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference disclosed, its performance illustrates the experiment’s central theme. More visible reasoning does not automatically mean better execution.

As an affiliate, we earn on qualifying purchases.
A public stress test, not a polished demonstration
Firmulate’s value is not that it presents an AI company as flawless. The attraction is the opposite: the company is publicly losing money, accumulating rules and exposing the distance between knowing and doing. Its synthetic employees can be followed through the company’s published quotes, while the live view keeps the financial countdown and daily activity visible.
For readers interested in solar and backup power, the broader lesson is straightforward. Reliability is demonstrated under strain, not inferred from a smooth demonstration. Firmulate brings that principle into business technology by testing whether an AI workforce reads deeply, protects trust and finishes the work it starts. The experiment’s most revealing moments are not spectacular failures. They are the nearly completed tasks—the correct diagnosis, the persuasive pitch and the missing signature—that separate apparent intelligence from an outcome a company can actually use.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI model monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model trustworthiness assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
