
AI and Automation
Pilots are easy. Outcomes are hard.
SUMMARY
Why AI implementations stall on the way to production, and what to look for in a partner before you sign.
AI pilots usually fail after launch, not before it, when production surfaces exceptions, organizational friction, and system gaps the demo never had to handle. The forward deployed engineer model exists to close that gap by owning the deployment from kickoff through the outcome.

Eui Chung
CTO
·
9 mins
Most product pilots don’t fail to launch. They launch fast, look impressive in week one, and generate genuine excitement. The failure, when it happens, shows up later, after the demo energy fades and the system has to hold up against real volume, real edge cases, and real organizational friction.
That’s the gap worth paying attention to: not “did the pilot go live,” but “did production ever deliver the outcome the business actually needed.” And the difference between those two outcomes usually isn’t the model. It’s the people who carried the project from one to the other.
If you run claims operations at a warranty company, you have probably watched this happen at least once. A vendor demos well, the pilot clears a clean subset of claims, and then the thing quietly never reaches the volume where it would have mattered. The instinct to be skeptical of the next one is earned, and it is part of why the build versus buy decision is so hard to get right.
Why do AI pilots fail to reach production?
There are two ways implementations commonly go wrong, and they look almost opposite on the surface.
The first is the slow pilot, a partner that insists on a perfect integration before ever testing against real data or real users. Momentum dies in planning meetings before anyone learns anything.
The second, more common failure is the fast pilot that stalls. A vendor ships a flashy demo quickly, stakeholders are impressed, and then production reveals everything the demo didn’t have to deal with: scale, exceptions, and the organizational adoption required to make the system stick. The pilot worked in the sense that it launched. It did not work in the sense that mattered.
Experienced, hands on implementation teams solve for both failure modes at once, moving quickly without sacrificing the rigor that production actually requires.
What does a well-run pilot look like?
Speed at the pilot stage is not about skipping steps. It is about not re-learning things that have already been learned elsewhere.
A team that has done this before recognizes common patterns on day one: which workflows tend to hide the most edge cases, which stakeholders need to be looped in early, and which risks are worth de-risking before they become blockers. A generic implementation team has to discover the shape of a customer’s business during the pilot itself. An experienced team already knows the shape of the problem in general, and spends the pilot adapting that knowledge to the specifics in front of them.
In warranty specifically, that means knowing before kickoff that the exception paths matter more than the happy path, that the parts data is going to be inconsistent, and that the adjusters who will actually use the system need to be in the room well before go-live.
That difference compounds. It is not just a faster kickoff, it is fewer false starts, fewer wrong turns, and a pilot that is actually representative of what production will require, rather than a stripped down best case scenario.
What separates a deployment that goes live from one that delivers?
“Went live” and “achieved the outcome” are not the same milestone, and it is worth being blunt about that. Plenty of product deployments launch successfully and never move the metric the business cared about in the first place: cost, cycle time, customer experience, whatever the original business case was built on.
In warranty claims, that means severity, cost to serve, and cycle time. If those three numbers look the same six months after go-live as they did before it, the deployment did not work, regardless of what the status report says.
Closing that gap typically comes down to three things:
Overcoming complexity. Production volume surfaces exceptions that never showed up in a pilot’s smaller, cleaner test set. Someone has to recognize these quickly and resolve them without derailing the broader rollout.
Navigating organizations. Alignment at kickoff is easy. Alignment six weeks into a messy rollout, when individual contributors are skeptical, IT has competing priorities, and operations wants results now, is the harder and more important version of the same problem.
Bridging systems. Legacy tools, inconsistent data formats, and manual handoffs were rarely designed with automation in mind. Getting them to function end to end, not just in a controlled test environment, is often the least glamorous and most decisive part of the work.
Here is the shape it usually takes in a claims environment. Auto-adjudication performs well in early testing against a few thousand claims. At real volume, a subset of cases starts breaking it: appliance model numbers formatted three different ways in a system nobody has migrated since 2014, or a coverage exception process that lives in an adjuster’s head and was never written down.
The fix is not a new model. It is someone with enough pattern recognition to spot the issue quickly, trace it to its source, and adjust the workflow without stalling the rest of the rollout. That kind of resolution comes from having solved a version of this problem before, not from reading about it for the first time.
Who owns the outcome after you sign?
The answer to that question is the best predictor of whether a deployment lands, and it is worth asking every vendor directly.
At ProPay the answer is a forward deployed engineer, or FDE. The FDE works inside your environment rather than from behind a product roadmap. They own the deployment end to end: adapting the software to the systems you actually run, resolving the exceptions that only appear at production volume, and working through the organizational friction that decides whether your team adopts the system. When something breaks at month three, they are the same person who was there at kickoff.
I have spent my career deploying hundreds of applications across large, complex organizations, working across dozens of business units, each with its own priorities, its own systems, and its own way of doing things. Some of the hardest problems I have had to solve had nothing to do with technology at all. They were about navigating internal bureaucracy, getting competing stakeholders to agree on a single path forward, and managing the egos and territorial instincts that show up whenever a new system threatens to change how people have always worked.
That experience is exactly why I believe so strongly in the FDE model. The technical build is rarely the hardest part of a deployment. The hardest part is everything around it: the politics, the legacy processes nobody wants to admit are held together with manual workarounds, and the stakeholder who is quietly skeptical because the last transformation project went nowhere.
You do not learn how to handle that from a playbook. You learn it by having done it, repeatedly, across enough different organizations to recognize the pattern the moment it shows up again. That is the difference between a team that can install software and a team that can actually get an organization to change how it operates.
Why speed and outcomes are not a trade-off
The deployment arc that actually works looks something like this: an accelerated pilot, a phased production rollout, active outcome tracking, and continued iteration based on real usage, not a hard stop the moment the system goes live.
The important insight is that speed and outcomes are not produced by two different teams or two different phases. The same real world pattern recognition that made the pilot fast is exactly what prevents production from stalling later. It is one continuous capability, applied consistently, rather than a handoff from an implementation team to a generic support queue once the contract is signed.
What this looks like in warranty claims management
Acceleration and outcomes are not competing goals. They are both the product of people who have solved these problems before and know how to apply that experience to a new environment, quickly and without guesswork.
At ProPay, this is what our forward deployed engineers bring to every implementation: the real world experience to navigate complex organizations, solve hard problems as they surface, and accelerate our customers toward the objectives and outcomes they set out to achieve.
Every home warranty carrier running ProPay is supported by a forward deployed engineer from first integration through production. Not as a professional services line item, but as the standard way our product gets deployed.
The reason we structured it that way comes back to the scorecard: severity, cost to serve, cycle time. Those numbers do not move because software went live. They move because someone stayed in the account long enough to find the exceptions quietly eating them, and refined the workflow around what they found. That is what an FDE is accountable for: outcomes. It is the only measure of a deployment that has ever mattered.
See what this looks like on your claims. We will run an analysis on ten thousand of your historical claims and show you where the outcome gap actually sits, stage by stage.
Frequently asked questions
Why do AI pilots fail to reach production?
Two ways, and they look opposite. Slow pilots die in planning, chasing a perfect integration before anyone tests against real data. Fast pilots stall after launch, when production surfaces the scale, exceptions, and adoption problems the demo never had to handle. The second is more common and harder to see coming.
How do you avoid a pilot that never goes into production?
Design the pilot to be representative rather than flattering. Test against real data and real users early, include the messy workflows instead of excluding them, and require that the people who run the pilot are the same people who will own production. A pilot that avoids the hard cases is not a smaller version of production, it is a different thing entirely.
What is a forward deployed engineer?
An engineer who works inside the customer's environment rather than from behind a product roadmap. They own the deployment end to end: adapting the software to the systems the customer actually runs, resolving exceptions that appear only at production volume, and working through the organizational friction that determines adoption.
Who owns the implementation if it stalls?
Worth asking every vendor directly, because the answers differ more than the pitches do. Some hand off to a support queue once the contract is signed. Under the forward deployed model the engineer who ran the integration is the same person accountable at month three, which is when most deployments actually get decided.
How long should a warranty claims automation pilot take?
Long enough to hit real exception volume and short enough that momentum survives it. The wrong question is how fast the pilot launches. The right one is whether the pilot was representative enough that production holds no surprises.

