MIKE NANTO
Case file 01 · full record · Warranty-claims platform

The same platform.
A third of the bill.
Built to grow.

A claims operation processing thousands of billed claims a month ran on a single hand-configured server nobody dared touch. Here is the full story of rebuilding it underneath live traffic, in plain terms.

~70%lower running cost, confirmed by the real bills after the move, trending to ~81%
0minutes of scheduled downtime at switchover, with no maintenance freeze
675automated tests guarding the code, run on every change
liveproblem detection: issues that once surfaced in ~2 weeks now surface in minutes
§1 · The situation

One server, holding up a business

The platform worked (thousands of warranty claims moved through it every month), but everything ran on one server that had been set up by hand over years. Nobody could rebuild it if it failed. Every deploy risked the whole system. Costs grew in lockstep with every new client. And when something broke, the operation often found out from a customer, weeks later.

The mandate: modernize it without stopping the line. Claims don't pause, and neither could the business.

§2 · What was done

Rebuilt in flight, cut over without a freeze

  • Re-architected the platform as six smaller systems that update independently, so one change no longer risks the whole business. Five weeks from first line of code to live.
  • Rebuilt every feature to match the old system, then ran the new system side by side with the old one on real work and compared the results before anyone depended on it, including the AI-assisted claim-submission process covering 13+ manufacturer and administrator programs.
  • Switched to the new system while the business kept running, with no maintenance freeze and no customer-facing outage window: a handful of minor same-day issues, all resolved.
  • Wrote the entire setup down as a repeatable recipe, with 675 automated tests checking every change. The platform can be rebuilt from scratch on command, and no deployment depends on someone remembering a password.
  • Stood up live error monitoring: problems that used to surface through customer complaints roughly two weeks later now surface in minutes, with the failing component named.

The discipline is the point: every risky step had a rehearsal, a rollback, and a verification. That is why the riskiest project this business ever ran produced its least eventful go-live.

§3 · What the client gained

In plain terms

Cost

A third of the bill

Same product, ~70% lower verified running cost, and cost now grows slower than clients are added, so the savings compound.

Speed

Faster every day

Quicker to load and work in, especially for the agents handling claims. Less waiting means more claims processed per person.

Visibility

Problems surface instantly

Live monitoring replaced find-out-from-a-customer. Issues announce themselves the moment they happen, with the cause attached.

Security

Better protected

Passwords and keys locked in a managed vault, access tied to individual people instead of shared passwords, and every defect found during the rebuild permanently closed, with automatic checks that keep it closed.

Growth

Built to grow

The platform can take on new clients and new claim programs without another big rebuild. Capacity grows piece by piece, only where it's needed.

Continuity

No single point of failure

The one irreplaceable hand-built server is gone. Everything is reproducible, documented, and survivable (including me leaving).

§4 · A note on names

This client, like every client here, is not named: confidentiality is part of what they pay for. Every figure comes from the engagement's own records, and a reference who'll take your call is available on request.

Running on a system nobody dares touch?

Thirty minutes is enough for me to tell you honestly whether it can be rebuilt in flight, and roughly what it would take.