Run to Failure: When Is Doing Less Maintenance the Rational Choice?
Run-to-failure is not automatically bad maintenance. It can be rational when failure is obvious, consequences are bounded, repair or replacement is quick, safety and compliance are not compromised, secondary damage is unlikely, and preventive or condition-based work would cost more than the risk it removes. The opposite is equally important: when failure is hidden, dangerous, production-critical, slow to recover from or capable of damaging other assets, waiting for failure is usually an expensive way to discover that the consequence mattered more than the component.
The Decision in One Sentence
Use run-to-failure only when the full consequence of failure is acceptably small and recoverable; otherwise prevent, detect, find, redesign or provide redundancy before the failure becomes the decision-maker.
The Decision to Make
The question is not whether preventive maintenance is good and breakdown maintenance is bad. The useful question is narrower: for this failure mode on this asset, should we spend resources before failure, or deliberately wait until failure occurs?
That choice can lead to time-based preventive work, condition-based monitoring, failure-finding for hidden protective functions, run-to-failure, redesign or redundancy. NASA's reliability-centered-maintenance guidance explicitly treats these as different possible outcomes rather than one universal hierarchy.
The maintenance policy should therefore follow the failure mode and its consequence, not habit, vendor calendar or technology fashion.
Why This Decision Matters
Too little maintenance can create unplanned downtime, quality losses, safety exposure and expensive emergency work. Too much maintenance can also destroy value through unnecessary parts, labor, shutdowns and maintenance-induced failures.
NASA's current facilities maintenance requirements describe reliability-centered maintenance as a search for the most effective mix of proactive and reactive maintenance. They explicitly state that run-to-failure can be effective when it is a conscious decision based on comparison of failure risk and cost with the cost of mitigating that risk.
NIST manufacturing research points in the other direction when consequences are material: among surveyed establishments using less than 50% reactive maintenance, heavier use of predictive maintenance was associated with less downtime and substantially lower defect rates. That is association, not a universal causal guarantee, but it reinforces the need to reserve reactive maintenance for the right failure modes.
Explain It Simply
Think about two light bulbs.
The first bulb lights a storage room. If it fails, someone notices immediately, replaces it in five minutes, nobody is endangered and nothing else is damaged. Replacing that bulb early every few months may cost more than simply keeping a spare and waiting for it to fail.
The second bulb is actually a warning lamp that tells operators a dangerous system has lost protection. If that lamp can fail silently, waiting until someone discovers the failure during an emergency is a completely different risk.
The component may look similar. The consequence is not.
Evidence Map
- Observed / NASA: reliability-centered maintenance seeks the most effective mix of proactive and reactive maintenance; run-to-failure is acceptable for some equipment when consciously selected through RCM analysis.
- Observed / NASA decision logic: maintenance choices should address functions, likely functional failures, consequences and what can reduce failure probability or consequence; if no suitable maintenance action exists and failure is unacceptable, redesign or redundancy becomes an outcome.
- Observed / NIST manufacturing study: establishments relying more on predictive maintenance within a less-reactive group were associated with 15% less downtime, 87% lower defect rates and smaller inventory increases related to unplanned maintenance. These are survey associations, not universal plant-level guarantees.
- Observed / ISO: ISO 55000:2024 and ISO 55001:2024 frame asset management around realizing value across the asset life cycle while balancing performance, risk and expenditure; the 2024 revision strengthens explicit decision-making requirements.
- Observed / OSHA: where equipment failure can create recognized serious hazards, preventive maintenance and verification of controls are part of hazard prevention; run-to-failure is not an acceptable shortcut around safety duties.
- Inference: run-to-failure is rational only when the consequence envelope is demonstrably small enough to carry.
- Uncertain: no public standard can choose the correct maintenance policy for a specific machine without its failure modes, duty, redundancy, repair time, cost, safety context and operating evidence.
The Real Options
- Run to failure. Do nothing before failure except keep the recovery path ready.
- Time- or cycle-based maintenance. Intervene at a defined interval when age or usage meaningfully changes failure risk and the task is technically effective.
- Condition-based maintenance. Monitor an indicator that gives enough warning to act before functional failure.
- Failure-finding. Periodically test hidden protective or standby functions whose failure may otherwise remain invisible.
- Redesign or redundancy. Change the system when maintenance cannot reduce an unacceptable consequence enough.
Five Tests Before Choosing Run to Failure
- Consequence: can failure harm people, product, environment, compliance, customers or other assets?
- Visibility: will the failure be obvious immediately, or can a protective function fail silently?
- Recovery: are spare, skills, access and restart procedures available fast enough?
- Secondary damage: can a small component failure damage a much more expensive system?
- Alternative effectiveness: is there a preventive or condition-based task that reliably changes the failure probability or provides useful warning at lower total cost?
Sidy's Synthesis — Maintenance Should Follow Consequence
Maintenance should not try to prevent every failure. It should prevent the failures whose price the system cannot afford.
My synthesis separates four questions: Consequence → Visibility → Recovery → Prevention value.
- Consequence: what is lost if this function disappears?
- Visibility: how quickly will we know?
- Recovery: how long and how much will it take to restore the function?
- Prevention value: is there an intervention that meaningfully changes the risk for less than the risk it removes?
This is not a named NASA, NIST or ISO framework. It is an analytical extension from their shared decision logic.
Decision rule: a failure is cheap only when its consequences are cheap too.
AI and Sensors Change the Cost of Knowing — Not the Consequence of Failure
Sensors, condition monitoring and machine-learning models can reduce the cost of detecting drift, combining signals and estimating remaining health. That can move some assets from time-based intervention toward condition-based maintenance.
But cheaper prediction does not automatically justify more sensors everywhere. A model can create false alarms, miss unobserved failure modes, drift when operating conditions change or optimize an indicator that is only loosely connected to the actual failure.
What becomes cheaper: data collection, anomaly detection and prioritization. What remains constrained: sensor quality, failure physics, access, spare parts, repair time and the physical consequence when the asset fails. Human judgment becomes more important in deciding which failure modes deserve monitoring, what threshold should trigger action and when the model itself is no longer trustworthy.
Second-Order Effects
A maintenance policy changes behavior. If every failure triggers an emergency, technicians learn firefighting rather than elimination of recurring causes. If every component is replaced early, teams can stop learning which failures were actually age-related. If every anomaly creates a work order, alert fatigue can make real warnings easier to miss.
Maintenance strategy therefore shapes the quality of future evidence. Record what failed, why, consequence, repair time, parts used and whether the chosen policy gave enough warning. The strategy should evolve when the evidence changes.
What Would Reopen the Decision
Reassess the maintenance policy when the asset becomes more critical, failure consequences rise, lead times or spare availability change, a hidden failure is discovered, operating duty changes, new monitoring becomes economically credible, repeated failures reveal a common cause, regulation or safety requirements change, or repair time begins to exceed the outage the business can absorb.
Remember This
The cheapest maintenance task is not always the cheapest maintenance strategy.
Primary sources
Facts, figures and quotations should be traceable to the sources below. Sidy's synthesis is labeled as synthesis and does not replace sourced facts.
- Facilities Maintenance and Operations Management — Chapter 7: Reliability Centered Maintenance — NASA
- Manufacturing Machinery Maintenance — National Institute of Standards and Technology
- ISO 55000:2024 — Asset management — Vocabulary, overview and principles — International Organization for Standardization
- ISO 55001:2024 — Asset management — Asset management system — Requirements — International Organization for Standardization
- Safety Management — Hazard Prevention and Control — Occupational Safety and Health Administration
