The honest math of predictive maintenance ROI
Predictive maintenance pays for itself when three numbers line up: failure cost, detection lead time, and alert precision. Here is how to run the math before you buy.
Start from the cost of failure, not the cost of sensors
Every credible PdM business case starts with a downtime ledger, not a hardware quote. Pull the last 24 months of unplanned failures on your candidate assets and price each one fully: repair labor and parts, expedited freight, overtime, production or service loss per hour multiplied by actual hours down, and any secondary damage. Most teams have never assembled this number, and most are surprised by it — a single failed compressor or chiller event frequently prices in the tens of thousands of dollars once production loss is counted.
Then classify which of those failures were detectable. Bearing wear, imbalance, misalignment, cavitation, insulation degradation, and fouling announce themselves through vibration, temperature, and current signatures — often weeks in advance. Sudden-death failures (a casting cracks, lightning strikes, an operator error) do not. PdM only earns money against the detectable class, so your addressable baseline is the detectable share of your failure ledger. Being honest about that share now prevents disappointment later.
This exercise also produces your target list for free. The assets that dominate the detectable-failure ledger — usually a short list of pumps, motors, compressors, air handlers, and conveyors — are exactly where instrumentation goes first. ROI comes from concentration, not coverage.
What the program actually costs
Count four cost lines. Hardware: sensors and gateways, plus installation labor — mounting a vibration sensor is fast, but permits and access on live equipment are not free. Software: the monitoring and CMMS capability, whether bundled or separate. Connectivity: usually minor with LoRaWAN or existing networks, occasionally material at remote sites. And attention: the human hours to review alerts, verify findings, and plan the resulting work. Attention is the line most business cases omit and the one that most often kills real programs.
The attention cost depends heavily on alert precision. A program generating high-quality, asset-contextualized alerts might need a few hours per week of a reliability engineer's time. A noisy program can consume a full headcount in triage and then get ignored — the industrial version of alarm fatigue. When comparing vendors, ask how alert thresholds are set and tuned, and what fraction of alerts customers act on. Vague answers here predict noisy programs.
Full-stack platforms shift this math by collapsing the integration and mapping costs that stitched sensor-plus-platform-plus-CMMS stacks carry. Whatever architecture you choose, insist that the quote covers installed-and-alerting, not shipped. The gap between those two states is where budgets die.
A worked example, stated as an illustration
Take an illustrative plant with 40 critical rotating assets. Suppose its 24-month ledger shows 18 unplanned failures on those assets averaging $25,000 fully loaded, of which two-thirds were vibration- or temperature-detectable. Addressable annual loss: roughly $150,000. Instrumenting the 40 assets with vibration-plus-temperature sensors, gateways, installation, and software might run $60,000–$90,000 in year one and a fraction of that annually after — figures vary widely by vendor and site, so treat these as placeholders for your own quote.
PdM does not prevent every detectable failure; it converts most of them from unplanned to planned. If the program catches 70 percent of detectable events and planned intervention costs a third of the unplanned equivalent, year-one savings land near $70,000–$80,000 against that $150,000 addressable base — roughly break-even in year one and strongly positive from year two, before counting the softer gains: less overtime, calmer scheduling, longer asset life from running equipment less often to failure.
The sensitivity analysis matters more than the point estimate. Run the numbers at 50 percent catch rate and at 85 percent; at your best and worst downtime cost estimates. If the case only works at optimistic values everywhere, narrow the asset list to the worst actors and re-run it. A concentrated program on 15 assets with a bulletproof case beats a broad program on 100 with a hopeful one.
Where PdM ROI goes to die
Failure mode one: alerts without ownership. A detection with no assigned human and no work order attached is a notification, and notifications get archived. The fix is workflow, not modeling — every alert above a severity line should create an inspection work order with an owner and a due date automatically. If your stack cannot do that in one system, that is a stack problem.
Failure mode two: no verified findings loop. Programs improve only when the outcome of each alert — confirmed fault, false alarm, watch — is recorded against the alert. Skip that discipline and thresholds never tune, precision never improves, and trust erodes quarter by quarter. Failure mode three: instrumenting by convenience rather than criticality. Sensors on easily accessible but unimportant equipment produce dashboards, not savings. Return to the failure ledger whenever the deployment plan drifts.
Finally, watch survivorship in your own reporting. The temptation is to count every caught anomaly as an avoided catastrophe at maximum cost. Auditors and CFOs see through it, and it corrodes the program's credibility. Count avoided events conservatively — planned-versus-unplanned cost delta on confirmed findings — and let the number be smaller and unimpeachable.
Sequencing a rollout that funds itself
Quarter one: instrument the top 10–15 worst actors, wire alerts to work orders, and assign a named owner for triage. Quarter two: tune thresholds against verified findings and publish the first catches internally — the planned bearing swap that pre-empted a weekend breakdown is your best internal marketing asset. Quarter three: expand to the next tranche using savings evidence from the first, and start trending energy and runtime data for efficiency findings, which often surface as a bonus.
By the end of year one you want three artifacts: a downtime ledger showing the trend on instrumented assets, a verified-findings log tying alerts to outcomes and costs, and a documented triage workflow that survives personnel changes. With those in hand, expansion stops being a proposal and becomes a budget line. That is what predictive maintenance ROI looks like when it is real: boring, documented, and compounding.