Skip to content
OpsSense
[ STRATEGY — 2026-02-24 — 8 MIN ]

Preventive vs reactive maintenance: the real economics

Run-to-failure is not always wrong and preventive is not always right. The economics turn on failure cost, failure pattern, and the price of intervention.

What run-to-failure actually costs

The visible cost of a reactive breakdown — parts and labor — is usually the smallest line. Industry rules of thumb put the true cost of an unplanned failure at three to five times the cost of the same repair done as planned work, once you count expedited freight, overtime, secondary damage, and production loss. A seized bearing replaced on schedule is a bearing. The same bearing run to seizure can take the shaft, the coupling, and a production shift with it.

Reactive-heavy operations also pay a hidden scheduling tax. When 60 percent or more of work arrives as emergencies, planners cannot plan. Wrench time drops because technicians spend their day being redirected. Parts inventory bloats because every stockout during a breakdown teaches the storeroom to over-buy. The maintenance backlog becomes a queue of interruptions rather than a managed portfolio.

The pattern is measurable in any CMMS with honest data: compare the average cost and duration of corrective work orders raised from breakdowns against corrective work raised from inspections. In most datasets the emergency version costs a multiple, not a margin. That ratio, computed from your own history, is the foundation of every maintenance-strategy business case.

When reactive is the rational choice

Run-to-failure is a legitimate strategy for the right assets. If a component is cheap, non-critical, quick to replace, and its failure causes no secondary damage — think general lighting, small fractional-horsepower fans, redundant utility pumps with automatic failover — then preventive attention is waste. The intervention costs more than the failure. The mistake is not having reactive assets; it is having them by accident rather than by decision.

The test is a simple criticality screen: what happens when this fails, how fast can we restore it, and does failure propagate? Assets that pass — low consequence, fast restoration, no propagation — go on a documented run-to-failure list with spares stocked. Documenting the decision matters. It converts 'we ignore these' into 'we decided these,' and it gives you a defensible answer when an auditor or a new plant manager asks why there is no PM on a class of equipment.

The over-maintenance trap

The opposite failure gets less attention: preventive programs that inspect and replace far too often. Time-based PMs inherited from OEM manuals are typically conservative — the manufacturer optimizes for warranty exposure, not your labor budget. Studies of PM programs routinely find a large share of tasks that have never once generated a corrective finding. Each of those tasks costs labor, parts, and — importantly — intrusion: every time a machine is opened, there is a chance of inducing the failure you were preventing.

Reliability engineering has a name for the underlying issue: most failure modes are not age-related. Classic actuarial studies of failure patterns found that only a minority of failure modes show a wear-out zone where time-based replacement helps; the majority fail randomly or early after intervention. Time-based PM does nothing for random failures except add cost and infant mortality. Condition monitoring — watching vibration, temperature, current draw — is the correct tool for those modes.

The practical move is PM optimization: for each PM task, ask what failure mode it addresses and check the work history for evidence it finds anything. Kill or extend the tasks with no findings. Shift condition-detectable modes to sensors or inspection routes. Teams that run this exercise typically cut PM labor meaningfully while coverage of real failure modes improves, because attention moves from calendar rituals to evidence.

Finding your mix

A defensible portfolio for a typical industrial site lands somewhere near: a small run-to-failure class chosen deliberately; time-based PM where failure genuinely is age- or usage-related (filters, belts, lubrication, calibration, statutory inspections); condition-based coverage on critical rotating and electrical equipment; and predictive models where telemetry history is deep enough to support them. The percentages vary by industry — a hospital and a cement plant should not have the same mix — but every asset should have an assigned strategy and a reason.

Sequence the transition by criticality and instrumentability. Start with the ten worst actors in your downtime log: these assets have both the economic case and the management attention. Put condition monitoring on them, tune out the noise, and let the first caught failure fund the next tranche. A visible early win — a bearing changed on a planned Tuesday instead of a chaotic Sunday — converts skeptics faster than any strategy deck.

Making the shift stick

Strategy changes fail in the schedule, not the analysis. Moving from reactive to planned work requires protecting planned hours from being cannibalized by emergencies — which feels impossible precisely when reactive load is high. The way through is gradual ring-fencing: reserve a fixed block of planned work per week, small at first, and grow it as breakdowns fall. Track schedule compliance weekly. It is the single best leading indicator that the transition is real.

Measure the shift with a small, stable dashboard: planned-versus-reactive work ratio, schedule compliance, PM compliance on critical assets, and downtime on the worst-actor list. Resist adding more metrics. When the reactive share drops below roughly a third of hours and stays there, the compounding begins — planners plan, parts arrive before jobs, and the same headcount produces visibly more uptime. That compounding, not any single repair, is where the economics of prevention actually live.

See your operations run themselves.

A 30-minute walkthrough with an operations engineer. Your assets, your workflows, your questions.