AI and the reliability engineer
Automated detection removes the data-gathering half of reliability engineering and makes the analytical half harder. The role moves from finding problems to deciding which are worth solving, and to proving the programme is working.
| Topic | Key point |
|---|---|
| The easy half got automated. The hard half did not. | Reliability engineering has always had two halves: gathering evidence about how assets fail, and deciding what to do about it. |
| Where the role moves | Then Now Collect vibration routes; build the dataset Data arrives; the work is deciding what it means for the schedule Argue for monitoring on critica |
| Deciding what not to monitor is now the higher-value call | The instinct when monitoring gets cheap is to instrument everything. |
| Proving the programme works is the part nobody plans for | The question that decides whether a programme keeps its budget is: did the findings change outcomes? |
| What analytics cannot decide | Consequence. A model ranks probability of failure. |
The easy half got automated. The hard half did not.
Reliability engineering has always had two halves: gathering evidence about how assets fail, and deciding what to do about it. Condition monitoring and analytics have largely taken the first.
That is a genuine gain, and it exposes how much of the second half was never really done. When evidence was expensive, arguing about priorities was cheap. When every critical asset produces a continuous stream of condition data, the constraint moves to deciding which findings justify acting on — and that is analysis, not detection.
Where the role moves
| Then | Now |
|---|---|
| Collect vibration routes; build the dataset | Data arrives; the work is deciding what it means for the schedule |
| Argue for monitoring on critical assets | Decide which assets should not be monitored, and defend that |
| Investigate failures after the fact | Prevent the ones with a usable P-F interval; still investigate the ones without |
| Set intervals from manuals and experience | Set them from observed degradation on your own assets |
| Report on work completed | Prove findings changed outcomes — much harder to measure |
Deciding what not to monitor is now the higher-value call
The instinct when monitoring gets cheap is to instrument everything. It is the wrong instinct, and resisting it is one of the clearest ways the role adds value.
Every monitored asset consumes analyst attention. A plant that monitors two hundred assets with capacity to properly interpret forty has not built a reliability programme; it has built an alarm queue. Criticality analysis — which assets justify attention because of what their failure costs — is what keeps the programme survivable, and ISO 17359 puts that audit before measurement selection for exactly this reason.
The defensible answer to "why is this pump not monitored" is a documented criticality decision, not an oversight.
Proving the programme works is the part nobody plans for
The question that decides whether a programme keeps its budget is: did the findings change outcomes? It is surprisingly hard to answer, because a prevented failure leaves no evidence.
What can be measured, if the recording is disciplined from the start:
- Findings that changed a schedule — not alerts raised. An alert nobody acted on prevented nothing.
- Confirmation rate — when the machine was opened, was the diagnosis right? This is the only honest measure of whether the analysis is trustworthy, and it needs the comparison to be recorded at the time.
- Ratio of planned to unplanned work on monitored assets, against the same ratio before.
- Failures that occurred anyway, split into modes the monitoring could see and modes it could not. The second group is a scoping problem; the first is an analysis problem, and conflating them hides both.
ISO 14224 gives a structure for the underlying failure and maintenance records. Without consistent recording none of the above can be computed later, which is why the measurement design has to exist before the programme starts rather than when the budget is questioned.
What analytics cannot decide
- Consequence. A model ranks probability of failure. What that failure costs in safety, environment, production and quality is an engineering and business judgement.
- Whether the failure mode matters. Detecting a mode that never caused a consequential failure is a solved problem nobody had.
- Design-out decisions. The best reliability intervention is often eliminating the failure mode, which no monitoring system will propose.
- Trade-offs against production. Taking an outage is a negotiation, not a calculation.
Frequently asked questions
Does AI replace reliability engineers?
It replaces the evidence-gathering half of the role and makes the judgement half more prominent. Deciding which assets deserve monitoring, what a finding is worth acting on, and whether to design the failure mode out entirely are all decisions that require consequence analysis a model does not perform.
How do you prove a predictive maintenance programme is working?
Measure findings that changed a schedule rather than alerts raised, and record whether the diagnosis was confirmed when the machine was opened. Track the planned-to-unplanned ratio on monitored assets, and split failures that happened anyway into modes the monitoring could see and modes it could not. This requires disciplined recording from day one; ISO 14224 provides the structure.
Should every critical asset be monitored?
No. Every monitored asset consumes analyst attention, and a programme that monitors more than it can interpret becomes an alarm queue nobody trusts. Criticality analysis comes first — ISO 17359 puts the equipment audit before measurement selection — and a documented decision not to monitor something is a legitimate output.
What is the most common reason these programmes fail?
Organisational rather than technical: alerts with no named owner and no authority to change the maintenance schedule. The technology keeps working while the programme stops mattering, and it is usually abandoned quietly a couple of years in.
Related guides
Setting up a condition monitoring programme
ISO 17359 sets out the general flow for condition monitoring: audit the assets, select measurements, establish a baseline, set alert criteria, then diagnose, prognose and act. Most failed programmes skip the criticality audit and start at sensor selection.
Predictive vs preventive maintenance
In EN 13306, the European maintenance-terminology standard, predictive maintenance is a sub-type of preventive maintenance, not its opposite. The real dividing line is what triggers the work: a fixed interval, or a measured and forecast condition. Whether prediction is possible at all depends on the asset's P-F interval.
Examples of predictive maintenance
Eight concrete cases — pump, motor, gearbox, fan, compressor, heat exchanger, steam trap and switchgear — each traced the same way: the failure mode, the technique that sees it, the signal that appears first, and the decision it should trigger.
Software that helps
AVEVA Predictive Analytics
Early-warning analytics for critical process and power assets.
Emerson AMS
Asset management and condition monitoring for process plants.
Cognite Data Fusion
Industrial DataOps and digital-twin foundation.