
Imagine handing an AI the keys to your repair shop’s customer list, service pipeline and playbook. Before it touches a real appointment or parts order, you would want to know what it does when a fleet customer threatens to leave, a competitor makes a move, or someone claiming to be the CEO asks it to bend the rules. Firmulate’s live experiment puts AI models through that kind of pressure, then offers businesses a way to run the test against their own operations.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company under pressure
Firmulate ran frontier models through the worst week of the same small software company. Each faced the same customers, crises and temptations; decisions were versioned and auditable. The point was to observe management under pressure, rather than judge a model by how polished its chat sounds.
The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The benchmark also treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
Recognizing trouble is only part of the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap was captured in the experiment’s verdict: “Same diagnosis, same pitch — no signature.” For a garage, the parallel is practical: an AI might correctly identify a customer-retention risk or recommend a high-value fleet contract, but still fail to carry the decision through.
The deal depended on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a familiar operational challenge: useful information can be present in a business’s records and still be missed when someone needs to act.
The pressure also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of restraint matters when an agent can see customer details, pricing or internal plans.
More diligence did not guarantee a better finish
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four participants. K3 also ran without an effort parameter, using the API default, while the other models ran at xhigh—a fairness detail to keep in mind when reading the ranking.
The company itself is live and watchable at firmulate.com. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
From watching to testing your own business
For automotive businesses, the interesting next step is not to assume that a model that performs well in a benchmark is ready to manage a live service operation. A repair group could test how an AI handles customer churn, a competitor’s offer, a public relations problem or pressure to disclose information—using its own business context before relying on it in day-to-day work.
Firmulate’s enterprise pilot uses a read-only export of a company’s business to run the wargame and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The live experiment is a demonstration; the pilot is the route from watching a company under pressure to examining your own.

Put the decisions under pressure first
The experiment’s central lesson is that identifying a problem does not guarantee follow-through. Models refused manipulation, but most failed to complete a deal their own analysis supported. An automotive business considering AI agents can test the decisions that matter to its customers and operations using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
