Supply Chain · Make
Maintenance Strategy
Match the policy to the asset: run-to-failure for cheap redundancies, preventive for wear items, predictive for critical machines.
- Time45 min
- FormatSmall group
- StageMake
Maintenance Strategy: what it is and why it works
A maintenance strategy assigns each asset the maintenance policy that fits its criticality and how it fails, instead of placing all equipment on the same calendar. Run-to-failure accepts breakdowns for cheap, non-critical or redundant items, provided spares are on the shelf and failure has no safety or environmental consequence. Preventive maintenance replaces or services components at fixed intervals where wear is age-related and predictable. Predictive or condition-based maintenance monitors indicators such as vibration, temperature, oil condition or motor current and intervenes when a developing failure is detected. The team lists critical assets and their failure modes, assigns a policy to each, and reviews the mix as conditions and costs change.
The logic comes from reliability-centered maintenance (RCM), whose foundation was laid by Nowlan and Heap in 1978. Their work showed that many failure modes are not related to age, so replacing parts on a fixed calendar does not prevent those failures and can even introduce new ones through intrusive work. For such modes, detecting the early warning signs of failure is more effective. Matching policy to failure mode avoids both wasted preventive work and costly unplanned breakdowns on critical machines. The strategy supports OEE improvement by reducing breakdowns on the constraint, sets spare parts buffers, and feeds stress scenarios for critical equipment. All maintenance work should follow site safety procedures, including lockout/tagout.
What you need
- An asset list with criticality: effect of failure on safety, environment, production and cost
- Known failure modes for critical assets, from history and technician experience
- Maintenance history: work orders, breakdowns, repair times and costs
- Available condition monitoring options and their cost
- Spare parts stock and lead times for key components
What you get
- A policy per asset or failure mode: run-to-failure, preventive or predictive
- A condition monitoring plan for critical assets
- Adjusted preventive maintenance intervals and removed low-value tasks
- Spare parts requirements linked to run-to-failure and critical assets
- An annual review schedule for the strategy
When to use it
When all equipment gets the same maintenance calendar regardless of risk.
How to do it, step by step
- List critical assets and their failure modes.
- Classify each: run-to-failure, preventive, predictive.
- Move critical machines toward condition monitoring.
- Standardize cheap redundancies to run-to-failure with spares on shelf.
- Review the mix yearly as sensor costs fall.
Worked example: Matching policies at a wastewater treatment plant
Illustrative scenario — figures are realistic but not from a real company.
A municipal wastewater treatment plant serving about 150,000 people maintained more than 800 assets on a calendar-based schedule. Technicians spent much of their time on routine preventive tasks, yet the plant had suffered two aeration blower failures in a year, each causing permit risk and costly emergency repairs.
- The maintenance team ranked assets by the consequence of failure. The large aeration blowers, influent pumps and UV disinfection system ranked highest; many small chemical dosing pumps and valves had installed spares.
- For the blowers, they listed failure modes such as bearing wear, imbalance and motor insulation breakdown, and found that most were not tied to operating hours.
- They moved blowers and influent pumps to condition monitoring with online vibration sensors and quarterly oil analysis, keeping preventive tasks only for filters and belts.
- Small dosing pumps with standby units were moved to run-to-failure with spare pumps on the shelf, and redundant preventive inspections were removed.
Result. Over the next year, vibration monitoring detected a developing bearing defect on one blower, which was repaired in a planned outage. Unplanned critical failures dropped, and technician hours on low-value preventive tasks fell by roughly a fifth. The team planned an annual review of policies as sensor costs and asset condition change.
Common pitfalls and how to avoid them
- Applying run-to-failure to items whose failure has safety or environmental consequences.Rule out run-to-failure where failure could harm people or the environment, regardless of the item's cost.
- Setting preventive intervals from manufacturer defaults without considering actual failure patterns.Use maintenance history and failure modes to adjust intervals, or switch to condition monitoring where failures are not age-related.
- Installing condition monitoring without defining who reviews the data and acts on alarms.Assign responsibility, alarm thresholds and response procedures before installing sensors.
- Choosing run-to-failure without stocking spares.Link each run-to-failure decision to spares on the shelf and a documented replacement procedure.
Frequently asked questions
What is the difference between preventive and predictive maintenance?
Preventive maintenance performs tasks at fixed intervals of time, cycles or operating hours, whatever the actual condition of the equipment. Predictive maintenance, also called condition-based maintenance, monitors indicators such as vibration, temperature or oil quality and intervenes only when signs of a developing failure appear. Preventive suits predictable wear; predictive suits failure modes that give a warning but are not tied to age.
When is run-to-failure maintenance acceptable?
It is acceptable when failure has no safety or environmental consequence, the item is cheap or redundant, the production impact is small or covered by standby equipment, and spares can be installed quickly. Light bulbs, small redundant pumps and some instruments are typical examples. The decision should be deliberate and documented, not a result of neglect.
What is reliability-centered maintenance?
Reliability-centered maintenance (RCM) is a structured method for deciding what maintenance each asset needs to keep performing its function in its operating context. It analyzes functions, functional failures, failure modes and their consequences, then selects the most suitable and cost-effective task for each mode. Its roots are in aviation, and the approach was formalized in the 1978 report by Nowlan and Heap.
Origin
Maintenance strategy — RCM lineage, Nowlan & Heap, 1978.
Related methods
- OEE Improvement LoopMeasure availability, performance and quality on the constraint machine; the biggest of the three losses is…
- Strategic BuffersPlace buffers where they buy resilience cheapest: critical components, long-lead items, single-source parts.
- Stress ScenariosPlay out three shocks — supplier failure, demand spike, logistics blackout — and rehearse the first 48 hours…
More in “Make”
Turn materials into products at the constraint’s rhythm.