From Fault Alerts to Engineering Decisions: How Agentic AI Is Changing Predictive Maintenance

Inzonex ResearchPublished Sources checked 12 September 20267 minute read

Maintenance software is moving from raising alerts to proposing a diagnosis: grouping separate faults, part failures and work orders into one recurring defect for an engineer to accept or reject. The method is sound and the decision stays with a person. What decides whether it works at your site is the quality of your failure records, not the model.

What was announced

On 9 September 2026 Veryon announced a new version of its Defect Analysis product, which now pulls aircraft faults and part failures into the detection of chronic defects. The change is in what the software is asked to produce. Instead of a flag that a measurement is abnormal, it proposes an answer to the question every maintenance engineer asks: whether this has happened before for the same reason.

Company claims, not independent benchmarks

Veryon describes combining a verified data set of chronic defects reviewed by maintenance experts, drawn from over a decade of outcomes and representing what it states is roughly 30% of the world's fleet, with each operator's own maintenance history. Its "Adaptive Clustering" is described as an agentic engine that learns from engineers' approve, split and merge decisions, tuned per aircraft make and model, and the company says the process stays under human control at every step.

The announcement cites customer results of up to a 75 per cent reduction in troubleshooting time for new technicians and up to a 23 per cent reduction in downtime costs. These are the vendor's figures for its own customers. They are not independent benchmarks, the conditions behind them are not published, and "up to" describes a best observed case rather than an expected one. Treat them as a reason to ask questions, not as a planning input.

One more boundary is worth stating plainly: this is aviation maintenance, where structured fault reporting is a regulatory habit. Most factories do not record failures to that standard, which matters more than any difference between the algorithms.

Alert, diagnosis and decision

Confusion about maintenance AI usually comes from collapsing three different things into the word "prediction". They have different inputs, different failure modes and different owners.

What each level consumes, what it produces, and who remains accountable
LevelInputOutputCharacteristic failureAccountable
AlertingSensor or fault-code threshold crossedA flag: something is out of rangeAlarm flood; the same fault raised repeatedly as if newOperator judgement
DiagnosisFault history, replaced parts, work orders, operating contextA proposed cause linking separate eventsConfident grouping of unrelated recordsEngineer accepts, splits or rejects
DecisionThe accepted diagnosis plus cost, risk and schedulingA work package with a defined scopeActing on a plausible but unverified causeNamed person with authority
The investigation path before and after automated clusteringThe upper row is the traditional path: an alert is raised, someone searches records manually, an experienced engineer recalls a similar case, and a diagnosis follows. The lower row assembles evidence from several sources, clusters repeated events, proposes a candidate chronic defect and ends at engineer review, which remains the decision point.What moves is the search, not the judgementBoth paths end with a person. The difference is how much of the evidence gathering happens before they are asked.Traditional: depends on who is on shift and what they rememberAssisted: the same decision, reached with assembled evidenceAlert raisedManual searchEngineer recallsDiagnosisMulti-sourceevidenceClusteringCandidate chronicdefectEngineer reviewThe risk does not disappear, it moves: an unreviewed cluster becomes a confident but unverified cause.
The decision point does not move; the evidence gathering does. Inzonex diagram of the two investigation paths.

Most sites have level one and call it predictive maintenance. The value being claimed now sits at level two, and its characteristic failure is not a missed fault. It is a confident grouping of events that share a word rather than a mechanism.

Why the taxonomy decides the outcome

This problem has a standard, and it predates the current wave of software by nearly three decades. ISO 14224 sets out how to collect and exchange reliability and maintenance data: equipment data including a taxonomy and attributes, failure data including cause and consequence, and maintenance data including the action taken, resources used and downtime. Its defined failure modes are intended to work as a common vocabulary across organisations. The standard is written for the petroleum, petrochemical and natural gas industries, so using its structure in other plants is an adaptation rather than compliance.

That vocabulary is exactly what a clustering system needs and what most maintenance histories lack. If one technician writes "seal leak", another writes "leaking from gland" and a third logs it against the wrong asset position, no model can reliably know these are the same failure of the same item. It will either miss the pattern or invent one.

The data layers a maintenance diagnosis depends onFive stacked layers, narrowest at the top. From the base upwards: free text written by technicians, structured fault codes, asset hierarchy, operating context and finally verified outcome. Most plants hold the lower layers and lack the upper ones, which is what limits automated clustering.Clustering can only be as good as the layer below itMost sites hold the bottom two layers well and the top three badly. That, not model choice, sets the ceiling.Verified outcomeDid the fix stop the recurrence?Operating contextLoad, duty, environment at the timeAsset hierarchyWhich physical item, at which positionStructured fault codesA controlled failure-mode vocabularyFree textWhat the technician wroteInzonex diagram. The data categories follow ISO 14224: equipment taxonomy, failure data and maintenance data.
Each upper layer is what makes the layer below interpretable. Inzonex diagram; the data categories follow the ISO 14224 structure.

The cheapest useful step

Before evaluating any product, take twenty recent work orders for one asset class and try to group them by failure mode yourself. Whatever stops you doing it — missing position identity, free-text causes, no record of what was replaced, no note of whether the fix held — is precisely what will stop the software.

Where clustering invents connections

Grouping repeated events is useful when the group shares a physical mechanism. It is misleading when the group shares something incidental. The common false links are predictable enough to check for directly.

  • Shared vocabulary, different fault. Records grouped because a phrase recurs, not because the same part failed the same way.
  • Shared author. One technician's wording style pulls unrelated events together.
  • Shared period. Events clustered from one shutdown, one bad batch of consumables or one commissioning error, which is a single event rather than a chronic defect.
  • Changed part identity. A superseded part number makes one recurring problem look like two, or two look like one.
  • Wrong level of the hierarchy. A defect attributed to a pump when it belongs to the coupling, or to a fleet when it belongs to one installation.

The acceptance test is the same one a reliability engineer already applies: state the mechanism that explains every member of the group, and predict what removing the cause should do. If the mechanism cannot be stated, the cluster is a hypothesis, not a finding.

What this means for manufacturers

  1. Decide which question you are buying an answer to

    Detecting an anomaly, explaining a recurrence and scoping a work package are three products. Vendors sell them as one.

  2. Audit the records before the software

    Check whether failures carry a controlled failure mode, a correct asset position, what was replaced and whether the repair held. That list is the ISO 14224 structure. Building it is the entry cost.

  3. Keep approve, split and merge as engineering acts

    Where a system learns from those decisions, the decisions become training data. Record who made them and on what basis, or you cannot audit the model later.

  4. Demand the mechanism with every cluster

    Accept a proposed chronic defect only with a stated physical cause. Log rejected clusters too; they are how you find out where the data misleads.

  5. Measure recurrence

    The outcome to measure is whether the same failure stops coming back. Count repeat events per asset before and after, over a period long enough to include the failure interval.

  6. Treat vendor percentages as questions

    Ask which sites, over what period, measured against what baseline, and what happened at the sites that did not improve.

Limits and open questions

  • No independent evaluation exists. The published improvement figures come from the vendor and cannot be checked against a stated method.
  • Sector transfer is unproven. Evidence from aviation fleets says little about a plant whose work orders are free text.
  • Feedback learning cuts both ways. A system tuned by engineers' decisions inherits their errors as well as their expertise, and drift is hard to detect without held-back cases.
  • Savings depend on what you did before. A site with disciplined defect elimination has less to gain than one starting from alarm floods, so no general payback figure is meaningful.
  • Record quality is the hidden cost. Nothing published states what it costs to bring maintenance data to the standard these methods assume.

Repeated failures have always been under-investigated, so the direction of this work is overdue. The constraint sits where it already sat: whether a plant records failures well enough for anyone, human or otherwise, to see the pattern.

Questions

What changes when maintenance AI becomes agentic?

The output changes. An alerting system tells you a measurement is abnormal. An agentic system assembles evidence from several records, proposes that a group of separate events is one recurring defect, and presents that to an engineer to accept, split or reject. The search for evidence moves into software; the judgement stays with the engineer.

Does this replace condition monitoring?

No. Condition monitoring still detects that something is changing. What is being added sits after detection: correlating repeated events, part replacements and work history to find whether the same fault keeps returning rather than treating each occurrence as new.

Why do maintenance records decide whether it works?

Clustering works on what it is given. If failures are recorded as free text with inconsistent naming, no controlled failure-mode vocabulary and an unclear asset hierarchy, the system will group records by wording rather than by physical cause. ISO 14224 exists because this problem predates AI by decades.

How do you tell a real chronic defect from a false link?

A chronic defect has a physical mechanism that explains every member of the group, and removing the cause stops recurrence. A false link is usually a shared word, a shared technician, a shared period or a shared part number that changed meaning. Require the mechanism before accepting the cluster.

Can these results be transferred to a factory from aviation?

The method transfers; the evidence does not. Aviation maintenance has strong incentives for structured fault reporting that most factories do not share. A plant with free-text work orders should expect weaker results than a fleet operator until its taxonomy improves.

Sources

Suggested citation: Inzonex Research (2026), From Fault Alerts to Engineering Decisions: How Agentic AI Is Changing Predictive Maintenance, published 12 September 2026. Product capabilities and improvement figures are the vendor's own statements. The three-level table, both diagrams, the false-link list and the acceptance test are Inzonex analysis.

Related industrial AI analysis