firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before the trip goes wrong

A storm cancels departures. A customer wants a refund. A competitor is ready to take the booking. For a travel company, the test of an AI assistant is not whether it can write a polished itinerary. It is what it does when several things go wrong at once—and whether it can turn its own advice into a sound decision.

Firmulate puts AI models through that kind of pressure in a watchable live experiment. Its latest test used a small software company, but the questions travel businesses face are familiar: Can an AI spot trouble, resist a dubious request and follow through when the stakes are real? For enterprises ready to move from watching to trying it themselves, Firmulate offers a pilot using a read-only export of their business.

One company, the same worst week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Every decision was versioned and auditable. The league ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline result was reassuring, then uncomfortable. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The diagnosis was right; the pitch was made. The signature never came. In a travel operation, that kind of gap could matter when a model identifies a way to retain a valuable customer or secure a booking, but stops short of acting on the recommendation.

The important clue was buried

The deal turned on a competitor weakness hidden two document references deep in the company’s own files—not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business context may sit in material an AI must actively consult, rather than in the immediate message or case in front of it.

That matters for travel firms handling booking terms, disruption policies, supplier agreements and customer histories. The experiment does not establish how a model would perform in those settings. It does show why a test built around a company’s own information and difficult scenarios can reveal more than a conversation demo.

Pressure came in through the side door

The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offered a more mixed lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, in a weaker form, in all four models. Detailed analysis alone did not guarantee disciplined follow-through.

There is a fairness caveat to the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also invites readers to try a “guess the model” quiz based on 242 real, unedited management decisions.

From watching to a company-specific test

Firmulate’s live company is a separate, watchable experiment: 13 synthetic employees operate with real money mechanics, burning €105k/month against €2.3k MRR. A public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday make the consequences visible. The league tests frontier models against the same company scenario; the live company shows ongoing activity. Both offer a view into how AI decisions can play out beyond a chat window.

For an enterprise, the next step is a pilot against its own business. Firmulate says the test can use a read-only export, run crisis scenarios against company-specific material and produce a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems. Travel businesses could use that setup to examine how models handle customer disruption, operational pressure and requests that appear to come from authority—before putting AI into workflows that affect staff or customers.

Watch Firmulate’s live experiment and see the league, quiz and ongoing company activity.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the hard week in the pilot

The league shows that spotting a crisis and refusing manipulation do not guarantee a completed business decision. A pilot gives a company a way to test those gaps against its own information and playbooks, while keeping the exercise read-only.

Travel and outdoor businesses interested in testing AI against their own worst-week scenarios can explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of AI Scalability: Focus On Plumbing, Not Just Model Advances

Conflicting 2026 adoption figures obscure a clearer signal: integration, governance and operating infrastructure are limiting AI agent deployments.

Epson Workforce Pro WF-4834 Wireless All-in-One Printer $100 + Free S&H

Epson’s Workforce Pro WF-4834 wireless all-in-one printer is available for $100 with free shipping, offering a cost-effective solution for home and small office use.

CUDA Books

A curated list of major CUDA programming books from beginner to advanced levels, including recent releases for 2024–2026, aims to support developers and researchers.

Translation Devices: Helpful Shortcut or Extra Gadget?

Understanding whether translation devices are essential tools or just gadgets depends on their limitations and proper use—discover more to make informed choices.