AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

If you run a solar installation business, you already know your worst week looks like: a flash sale of competitor panels undercuts your quotes, a utility changes its interconnection rules mid-quarter, a customer with a half-installed battery system threatens to go public, and someone impersonating you emails your ops team asking for an “urgent” payment redirect. Now imagine handing that week to an AI — and watching, safely, whether it signs the deal, spots the trap, or freezes.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is essentially what Firmulate has been doing in public. The project runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. The results are worth attention from anyone who will soon have AI agents touching their CRM, their install scheduling queue, or their sales forecast. Which, in home energy, is sooner than most people think.

The experiment

Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, like a commit log for management.

The final league table from July 2026 tells a story chat demos never show:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

For context, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, no amount of good work outweighs a breach of trust.

Same diagnosis, same pitch — no signature

Here is the finding that matters most for anyone evaluating AI for their business. All models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Think of it like an installer who correctly diagnoses that a customer’s rooftop needs a specific panel configuration, delivers a flawless proposal, and then never asks for the deposit. The analysis was perfect. The business outcome was zero.

And the buried fact is worse: the decisive competitor weakness — the thing that would have won the deal at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer call at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation won. The ones that skimmed didn’t. For solar companies sitting on years of job files, rebate paperwork, and interconnection records, that lesson lands hard: an AI that doesn’t mine your own history is leaving deals on the roof.

The manipulation test

Every model faced social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was strikingly clinical: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’ve ever worried about an AI assistant approving a wire transfer because the email looked like it came from you, this is the encouraging data point.

Thoroughness isn’t everything

The most surprising profile was Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Diligence without follow-through is a familiar failure mode in human businesses too.

One fairness note the project discloses: Kimi K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still placed second.

It’s live, right now

The experiment is watchable, not hypothetical. A live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k MRR — runs every business day with a public cash countdown and over 680 self-learned playbook rules, all versioned. And there’s a twist for the curious: 242 real, unedited management decisions power a “guess the model” quiz, so you can test whether you can tell an AI CEO from its rivals by decisions alone.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The gap the Firmulate experiment exposes is the one chat demos hide: models that diagnose brilliantly but don’t close, that read the customer but not their own files, that are honest but not always disciplined. You only find these failure modes by running the wargame — not by reading a spec sheet.

That’s exactly what the pilot offers. Enterprises can run the same wargame against a read-only export of their own business: their customers, pipeline, and rules, facing churn waves, price increases, competitor attacks, and social-engineering pressure. You get a board report with a model ranking and the weak points of your own playbooks — and nothing ever writes back to real systems. For a solar or home-energy business, that means you can watch an AI handle your worst week — a rebate cliff, a competitor undercutting your battery quotes, a fake-CEO payment scam — before any agent touches a live customer record.

Ready to stress-test your own playbooks? Explore the enterprise pilot at firmulate.com/pilot.html or reach out at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hawaii University Nearly 100% Solar Powered

Brigham Young University-Hawaii is advancing its solar project to become nearly fully powered by solar energy, with phase two underway and additional systems planned.

Rune Raises $40 Million Series A To Build On-site, Off-grid AI Data Centers At Solar Facilities

Rune raises $40 million in Series A funding to develop on-site, off-grid AI data centers powered by solar energy, aiming to enhance sustainability and data security.

Lenovo Surges In Global Coverage

Lenovo has experienced a notable surge in global media mentions, with GDELT reporting 23 mentions within a recent window, indicating increased international attention.

Wiring Solar Arrays in Series Vs Parallel

I’m here to help you understand whether wiring solar arrays in series or parallel best suits your needs, but the key differences might surprise you.