AI and the reliability engineer

Automated detection removes the data-gathering half of reliability engineering and makes the analytical half harder. The role moves from finding problems to deciding which are worth solving, and to proving the programme is working.

Key takeaways at a glance
TopicKey point
The easy half got automated. The hard half did not.Reliability engineering has always had two halves: gathering evidence about how assets fail, and deciding what to do about it.
Where the role movesThen Now Collect vibration routes; build the dataset Data arrives; the work is deciding what it means for the schedule Argue for monitoring on critica
Deciding what not to monitor is now the higher-value callThe instinct when monitoring gets cheap is to instrument everything.
Proving the programme works is the part nobody plans forThe question that decides whether a programme keeps its budget is: did the findings change outcomes?
What analytics cannot decideConsequence. A model ranks probability of failure.

The easy half got automated. The hard half did not.

Reliability engineering has always had two halves: gathering evidence about how assets fail, and deciding what to do about it. Condition monitoring and analytics have largely taken the first.

That is a genuine gain, and it exposes how much of the second half was never really done. When evidence was expensive, arguing about priorities was cheap. When every critical asset produces a continuous stream of condition data, the constraint moves to deciding which findings justify acting on — and that is analysis, not detection.

Where the role moves

ThenNow
Collect vibration routes; build the datasetData arrives; the work is deciding what it means for the schedule
Argue for monitoring on critical assetsDecide which assets should not be monitored, and defend that
Investigate failures after the factPrevent the ones with a usable P-F interval; still investigate the ones without
Set intervals from manuals and experienceSet them from observed degradation on your own assets
Report on work completedProve findings changed outcomes — much harder to measure

Deciding what not to monitor is now the higher-value call

The instinct when monitoring gets cheap is to instrument everything. It is the wrong instinct, and resisting it is one of the clearest ways the role adds value.

Every monitored asset consumes analyst attention. A plant that monitors two hundred assets with capacity to properly interpret forty has not built a reliability programme; it has built an alarm queue. Criticality analysis — which assets justify attention because of what their failure costs — is what keeps the programme survivable, and ISO 17359 puts that audit before measurement selection for exactly this reason.

The defensible answer to "why is this pump not monitored" is a documented criticality decision, not an oversight.

Proving the programme works is the part nobody plans for

The question that decides whether a programme keeps its budget is: did the findings change outcomes? It is surprisingly hard to answer, because a prevented failure leaves no evidence.

What can be measured, if the recording is disciplined from the start:

  • Findings that changed a schedule — not alerts raised. An alert nobody acted on prevented nothing.
  • Confirmation rate — when the machine was opened, was the diagnosis right? This is the only honest measure of whether the analysis is trustworthy, and it needs the comparison to be recorded at the time.
  • Ratio of planned to unplanned work on monitored assets, against the same ratio before.
  • Failures that occurred anyway, split into modes the monitoring could see and modes it could not. The second group is a scoping problem; the first is an analysis problem, and conflating them hides both.

ISO 14224 gives a structure for the underlying failure and maintenance records. Without consistent recording none of the above can be computed later, which is why the measurement design has to exist before the programme starts rather than when the budget is questioned.

What analytics cannot decide

  • Consequence. A model ranks probability of failure. What that failure costs in safety, environment, production and quality is an engineering and business judgement.
  • Whether the failure mode matters. Detecting a mode that never caused a consequential failure is a solved problem nobody had.
  • Design-out decisions. The best reliability intervention is often eliminating the failure mode, which no monitoring system will propose.
  • Trade-offs against production. Taking an outage is a negotiation, not a calculation.

Frequently asked questions

Does AI replace reliability engineers?

It replaces the evidence-gathering half of the role and makes the judgement half more prominent. Deciding which assets deserve monitoring, what a finding is worth acting on, and whether to design the failure mode out entirely are all decisions that require consequence analysis a model does not perform.

How do you prove a predictive maintenance programme is working?

Measure findings that changed a schedule rather than alerts raised, and record whether the diagnosis was confirmed when the machine was opened. Track the planned-to-unplanned ratio on monitored assets, and split failures that happened anyway into modes the monitoring could see and modes it could not. This requires disciplined recording from day one; ISO 14224 provides the structure.

Should every critical asset be monitored?

No. Every monitored asset consumes analyst attention, and a programme that monitors more than it can interpret becomes an alarm queue nobody trusts. Criticality analysis comes first — ISO 17359 puts the equipment audit before measurement selection — and a documented decision not to monitor something is a legitimate output.

What is the most common reason these programmes fail?

Organisational rather than technical: alerts with no named owner and no authority to change the maintenance schedule. The technology keeps working while the programme stops mattering, and it is usually abandoned quietly a couple of years in.

Related guides

Software that helps