Reliability Fundamentals: RCM, FMEA/FMECA & Choosing the Right Task Type
🎯 Who this is for: Reliability engineers and maintenance planners who will be making the actual maintenance decisions in MAS 9 — and want the decision logic straight before they translate it into Failure Codes, PMs, and condition monitoring.
Series: Part 1 of 7 — MAS 9 Reliability Implementation Playbook | Read time: 20 minutes
🎯 Why the Method Comes Before the Application
There is a temptation, on a MAS 9 project, to open Maximo and start building. Resist it for one part. A maintenance task configured against the wrong decision is not a small mistake you can tune away later — it is a recurring cost baked into a PM record that fires forever, consuming labour, parts, and downtime while doing nothing measurable to failure risk. The applications in this series are excellent at executing decisions. They are indifferent to whether the decision was right. That is your job, and this part is where you learn to do it.
The whole of reliability engineering can be compressed into a single objective:
Sustain the required function of the asset at the lowest sustainable total cost — where total cost = the consequence cost of failure + the cost of the maintenance effort spent to avoid it.
Read that carefully, because it contains the trap that sinks most programs. More preventive maintenance is not inherently better. Every task consumes labour, parts, and downtime, and beyond a point it adds cost without reducing risk — and, as Part 2 will show, intrusive maintenance can even cause failures through reassembly errors and contamination. A good strategy is therefore risk-based: it concentrates effort where failure consequences are severe and is deliberately minimal where they are trivial.
<aside>
💡 Key insight: The value of a maintenance task is avoided failure cost minus the cost of the task. If that number is negative, the task is destroying value even if it runs flawlessly. This is the single lens through which every decision in this part is made.
</aside>
Reliability Centered Maintenance (RCM) is the structured method for reaching those decisions defensibly, and Failure Mode and Effects Analysis (FMEA) is the analytical technique that feeds it. The rest of this part walks both, then hands you a decision matrix that maps any failure mode to exactly one maintenance task type — the output you will carry into Part 4 when you build it in Maximo.
🧠 RCM: The Seven Questions of SAE JA1011
RCM is not a software feature and it is not owned by any vendor. It is a discipline, and the international standard SAE JA1011 defines what a process has to do to earn the name RCM: it must answer seven questions, in order, for each function of the asset in its operating context.
- What are the functions and associated desired performance standards of the asset in its present operating context?
- In what ways can it fail to fulfil its functions? (functional failures)
- What causes each functional failure? (failure modes)
- What happens when each failure occurs? (failure effects)
- In what way does each failure matter? (failure consequences)
- What can be done to predict or prevent each failure? (proactive tasks)
- What should be done if a suitable proactive task cannot be found? (default actions)
Why "in its operating context" is not a throwaway phrase
The same pump has a different reliability strategy on a fire-water standby line than on a continuous cooling duty, even if the nameplate is identical. On the standby line its dominant risk is a hidden failure — it will not be called on until there is a fire, and by then it is too late to discover the impeller has seized. On the cooling duty the same seizure is immediately obvious because the process trips. Context changes the consequence, and consequence changes the decision. RCM forces you to answer question 1 for this asset, in this service — which is exactly why you cannot buy a generic PM library and call it a reliability strategy.
The proactive-task test
Question 6 has a two-part gate that most amateur analyses skip. A proactive task is worth putting into your program only if it is:
- Technically feasible — the task can actually detect or prevent the failure mode (an oil analysis will not catch a fatigue crack), and
- Worth doing — the task's cost is less than the consequence it mitigates over its lifetime.
If no task passes both tests, you fall through to question 7, and the default action depends on the consequence category — which is the heart of the method.
🔺 The Consequence Categories That Actually Drive the Decision
Here is the part that separates RCM from "we should PM everything." The decision about what to do is driven by how the failure matters, not by how old the asset is. RCM sorts every failure into one of four consequence categories, and the category dictates the default action when no proactive task qualifies.
| Consequence category | What it means | Default action if no proactive task qualifies |
|---|---|---|
| Hidden | The failure is not evident to operators under normal conditions (standby, protective, and backup devices) | Failure-finding — a scheduled test to reveal the hidden state; redesign if the failure-finding interval is impractical |
| Safety / environmental | The failure could hurt someone or breach an environmental limit | Redesign is mandatory — you may not leave a serious safety/environmental failure unmanaged |
| Operational | The failure costs production, throughput, or quality | Redesign is justified only if the economics work; otherwise accept and manage |
| Non-operational | The failure costs only the direct repair — no safety, environmental, or production impact | Run-to-failure is acceptable |
<aside>
💡 Key insight: "Run-to-failure" is a legitimate, engineered decision — not negligence — when the failure has only minor economic consequences and prevention would cost more than the failure itself. The mistake is not choosing run-to-failure; the mistake is choosing it by default because nobody analyzed the mode.
</aside>
The hidden-failure category is the one operations teams under-serve most, because by definition nothing tells you the device has failed. A pressure relief valve that has silently corroded shut gives no warning until the overpressure event it was meant to protect against — at which point you have two failures at once. That is why failure-finding tasks exist: their entire purpose is to reveal hidden failures on a schedule, before the demand arrives.
🔍 FMEA and FMECA: Finding and Ranking the Modes
RCM tells you what questions to ask. FMEA is the technique that produces the raw material for questions 2 through 5.
- FMEA (Failure Mode and Effects Analysis) works bottom-up: for each component, it enumerates every way it can fail (the failure mode), what causes that mode, and what effects the mode produces locally and on the wider system.
- FMECA (…and Criticality Analysis) adds a criticality ranking so you can prioritize the modes rather than treating them as a flat list.
The RPN screen, worked
The most common quantitative screen in FMECA is the Risk Priority Number:
RPN = Severity × Occurrence × Detection — each scored on a 1–10 scale.
- Severity (S): how bad the effect is if the mode occurs (10 = catastrophic/safety; 1 = trivial).
- Occurrence (O): how likely the mode is (10 = almost certain; 1 = extremely unlikely).
- Detection (D): how hard it is to detect before it causes harm — note the inversion: 10 = you will not catch it in time; 1 = you will catch it easily.
Take a mechanical seal failure on a critical process pump:
| Factor | Score | Reasoning |
|---|---|---|
| Severity | 8 | Seal failure trips the process and spills product — high impact, no injury |
| Occurrence | 6 | Seals are a known recurrent failure on this duty |
| Detection | 4 | Seal-pot level and leakage give a reasonable early warning |
| RPN | 8 × 6 × 4 = 192 | High enough to demand a proactive task |
Now compare a bearing on the same pump scored S=7, O=5, D=3 → RPN = 105. The seal mode ranks above the bearing mode, so it gets attention first. That is the entire point of the number: it orders your queue.
Where RPN misleads you
<aside>
⚠️ Watch out: RPN is a screening aid, not an absolute measure of risk. A mode with Severity 10 (a safety failure) can warrant mandatory action even if its overall RPN is modest, because a low occurrence or good detection score arithmetically dilutes a hazard you are not allowed to tolerate. Always let severity override the multiplication for safety and environmental modes.
</aside>
Two other RPN traps worth naming: the scale is ordinal, not ratio (an RPN of 200 is not "twice as risky" as 100), and the number is gameable — an analyst under pressure to hit a threshold can quietly nudge a Detection score. Treat RPN as a conversation-starter for the engineering judgement in questions 5–7, not a replacement for it.
🛠️ Choosing the Right Task Type
This is the output that matters. For every failure mode you keep, you choose exactly one maintenance task type. The choice is driven by the nature of the mode and its consequence, not by habit.
| Task type | Choose it when… | Detects/prevents by |
|---|---|---|
| Run-to-failure (RTF) | Consequences are minor and prevention costs more than the failure | Nothing — you plan the repair, not the prevention |
| Time / usage-based PM | The failure mode is genuinely age-related (a definable wear-out or safe-life point) | Restoring or replacing before the wear-out point |
| Condition-based (CBM) | A measurable degradation signal exists with useful lead time | Inspecting a parameter and acting on a threshold |
| Predictive (PdM) | Lead time and asset value justify trend modelling / analytics | Modelling the trend to forecast the failure point |
| Failure-finding (FF) | The failure is hidden (standby / protective devices) | Scheduled testing to reveal a failed state |
| Redesign | No task adequately controls a serious failure mode | Eliminating or de-rating the mode itself |
Worked selections
- Pump mechanical seal (process duty): measurable signal (seal-pot level, leak detection) with lead time → CBM, backed by a job plan that triggers on the action limit.
- Fire-water standby pump seizure: hidden — nothing reveals it until demand → failure-finding (a scheduled monthly run-test).
- Motor bearing on a non-critical fan: minor consequence, cheap to replace, no injury or production hit → run-to-failure, with spares stocked.
- Gearbox with a defined design life and no in-service degradation signal: genuinely age-related → time-based PM at the safe-life interval.
- High-value turbine with sensors and a fleet history: value and lead time justify analytics → predictive, using Maximo Predict (Part 5) on top of the CBM signal.
<aside>
💡 Key insight: Notice how few modes land in "time-based PM." That is not an accident of these examples — it is the central finding of reliability science, which Part 2 quantifies: most failure modes are not age-related, so a calendar-based PM cannot reduce their probability. When in doubt, a mode is more likely CBM, RTF, or failure-finding than time-based PM.
</aside>
📉 The P-F Curve and Why It Sets Inspection Frequency
Condition-based and predictive maintenance both rely on one idea: most failures announce themselves before they finish. The P-F curve describes that announcement.
condition
▲
│● ─────────────╮ P = potential failure first detectable
│ ╲ (e.g. vibration rises, oil metal count climbs)
│ ╲
│ ╲ ← the P-F interval
│ ╲
│ ● F = functional failure
│ (the asset can no longer do its job)
└───────────────────────────▶ timeThe P-F interval is the time from the moment a degradation signal first becomes detectable (P) to the point of functional failure (F). It is the actionable warning window, and it dictates one thing above all:
Your inspection interval must be meaningfully shorter than the P-F interval. If you inspect once a month but the P-F interval is three weeks, you will routinely walk up to an asset that has already failed.
Worked P-F example
Suppose vibration analysis on a pump bearing gives a typical P-F interval of about eight weeks — that is, once the bearing defect frequency becomes detectable, you generally have roughly eight weeks before the bearing is functionally gone. A common rule of thumb is to inspect at half the P-F interval to guarantee at least two chances to catch it, so you set a four-week vibration route. Now the reliability logic connects to Maximo mechanics: that four-week interval becomes a meter-based or time-based condition-monitoring route (Part 3), and the action limit that trips a work order is the vibration level associated with point P.
Predictive maintenance is, in this framing, an effort to detect P earlier and estimate the slope: sensors and machine-learning models (Part 5) push the detectable point further left and turn "it will fail sometime in the window" into "it will fail in about 120 operating hours."
🧮 A Fully Worked FMEA: One Centrifugal Pump
Let us put the whole method on one asset — a critical centrifugal process pump, so you can see the seven questions and the task-type decision flow end to end.
Q1 — Function & performance standard: Deliver 200 m³/h of cooling water at 4 bar to the reactor jacket, continuously.
Q2–Q5 — the FMEA table:
| Failure mode | Cause | Effect | Consequence category | RPN (S×O×D) |
|---|---|---|---|---|
| Mechanical seal leak | Seal-face wear / dry run | Product leak, process trip | Operational + environmental | 8×6×4 = 192 |
| Bearing failure | Fatigue / contamination | Vibration, secondary seal damage | Operational | 7×5×3 = 105 |
| Impeller erosion | Cavitation / abrasive fluid | Gradual head loss below spec | Operational | 6×4×5 = 120 |
| Coupling failure | Misalignment | Immediate stoppage | Operational | 7×3×4 = 84 |
| Motor winding burnout | Insulation age / overload | Total loss of function | Operational | 8×2×6 = 96 |
Q6–Q7 — task selection per mode:
| Failure mode | Task-type decision | Why |
|---|---|---|
| Mechanical seal leak | CBM — seal-pot level + leak detection, action-limit WO | Measurable signal with lead time; highest RPN |
| Bearing failure | CBM — vibration route at half the P-F interval | Classic detectable degradation |
| Impeller erosion | CBM — periodic performance test (head vs flow) | Slow, measurable head loss |
| Coupling failure | Time-based PM — laser alignment check + inspection | No live signal; alignment drifts with runtime |
| Motor winding burnout | PdM / RTF — thermography + spare motor | Low occurrence; predictive if sensored, else run-to-failure with a spare |
Notice the outcome: of five modes, three are CBM, one is time-based, and one is predictive-or-run-to-failure. Not a single mode defaults to "monthly PM because that is what we have always done." That is what a real analysis produces — and it is the exact set of decisions you will encode into Failure Codes, condition-monitoring points, job plans, and PMs in Parts 3 and 4.
⚠️ Edge Cases and Common Mistakes
- Multiple modes on one component. A pump does not have "a failure" — it has a portfolio of modes, each needing its own decision. Teams that assign one PM "for the pump" have skipped the analysis entirely.
- Hidden failures hiding in plain sight. Backup pumps, relief valves, trip systems, and standby generators are the modes most likely to be under-maintained precisely because their failures are invisible. If you have no failure-finding tasks in your program, you have almost certainly missed hidden failures.
- RPN worship. Ranking by RPN and then acting only on the top decile silently ignores high-severity, low-occurrence safety modes. Screen with RPN; decide with consequence category.
- Context blindness. Copying a PM from a similar asset on a different duty imports the wrong operating context and often the wrong task type.
🚫 The Anti-Patterns to Retire Now
- PM-volume thinking. "We added 400 PMs this year" is a cost metric masquerading as a reliability metric. The right metric is per-failure-mode coverage with the right task type.
- Calendar worship. Time-based PM is the correct answer for genuinely age-related modes and roughly no others. Defaulting everything to a calendar interval is the most expensive habit in maintenance.
- Analysis paralysis. RCM done to textbook exhaustiveness on every bolt never finishes. Concentrate the full seven-question treatment on critical assets and use a lighter FMEA — or the vendor library in the Reliability Strategies app (Part 4) — for the rest.
📋 Practical Notes: Running Your First FMEA Workshop
- Pick one critical asset class (pumps, compressors, or a critical line) — not the whole plant. A focused first pass that finishes beats a plant-wide effort that stalls.
- Get the operators in the room. They know the failure modes and effects that never made it into any document. The maintenance planner alone will miss context.
- Timebox each function. Answer the seven questions, decide the task type, and move on. You are producing decisions, not a novel.
- Record the *decision*, not just the task. For every mode, note the chosen task type and the reason (consequence category, RPN, P-F interval). Part 3 turns those decisions into Failure Codes; Part 7's feedback loop revisits the reasons.
- Leave with a task-type column filled in. The deliverable of the workshop is exactly the two worked tables above — modes on the left, task-type decisions on the right. That is your build backlog for Maximo.
<aside>
💡 Key insight: Everything downstream in this series — the failure hierarchy, criticality, condition monitoring, job plans, PMs — is just the durable, auditable recording of the decisions you make in this workshop. Get the method right here and the applications become data entry. Get it wrong here and no amount of software fixes it.
</aside>
Key Takeaways
- RCM (SAE JA1011) is seven questions per function — and a proactive task earns a place only if it is both technically feasible and worth doing.
- Consequence category drives the decision, not asset age: hidden, safety/environmental, operational, or non-operational each has a different default action.
- FMEA finds the modes; FMECA ranks them. RPN = Severity × Occurrence × Detection is a screen to order your queue, never an absolute risk measure — let severity override for safety modes.
- Every failure mode maps to exactly one task type: run-to-failure, time-based PM, condition-based, predictive, failure-finding, or redesign — and most modes are not time-based PM.
- The P-F interval sets inspection frequency: inspect at meaningfully less than the P-F interval, and predictive maintenance is the effort to detect P earlier and forecast F.
References
IBM Official
- Maximo Manage — Reliability Strategies module (IBM Documentation)
- Maximo Manage — Failure Codes and failure reporting (IBM Documentation)
Standards & Community
- SAE JA1011 — Evaluation Criteria for RCM Processes
- Maximo Secrets — Failure Codes and Failure Reporting
- Maximo Secrets — Reliability Strategies
Series Navigation
| Previous: | Series Index — MAS 9 Reliability Implementation Playbook |
|---|---|
| Next: | Part 2 — Reliability Metrics That Matter: MTBF, MTTR, Availability & PM Optimization |
About TheMaximoGuys: We help Maximo teams navigate the move to MAS 9 with practical, no-hype guidance grounded in how the platform actually behaves — from architecture and migration planning to the day-to-day work of configuring, extending, and running Maximo.
Published by TheMaximoGuys | July 2026



