SPECIAL SECTION | Artificial intelligence

Where Artificial Intelligence Belongs in Vibration Analysis

The most useful question is not whether AI can diagnose a machine, but which part of the diagnosis it should be trusted with.

Oumayma Jerbi | BETAVIB

One question that always seems to come up with reliability engineers is: “Can AI actually read a vibration spectrum yet?”

And honestly, it is a good question. But a simple yes or no does not really answer it. AI has gotten good at handling technical information, so it is easy to think: Why not give it a spectrum, tell it a bit about the machine, add some trend data and see what diagnosis it comes back with?

It sounds like it should work. But vibration analysis is not quite that simple. After looking at what actually goes into a diagnosis, it is easier to see where the problems are.

Why a Language Model Cannot Read a Spectrum

Start with the direct approach: Feed the spectrum to the model and ask what is wrong. Four problems appear immediately, and none of them are solved by waiting for a larger model.

  1. Numbers are not quantities: A language model does not process 112.4 as a magnitude. It processes it as a short sequence of text tokens, in much the same way it processes a word. Asking such a system to reliably distinguish a ball pass frequency outer race component at 112.4 hertz from a gear mesh sideband at 114.8 hertz—across several thousand spectral lines—is asking it to perform arithmetic that its architecture does not do. It will produce an answer. The answer will often be confidently wrong, and it will frequently describea harmonic series that is not present in the data.
  2. The arithmetic of context: A single 3,200-line spectrum, written out as text, consumes roughly 40,000 to 50,000 tokens of a model’s working memory. A pump and motor set with 12 measurement points, each captured in velocity, acceleration and envelope domains, produces more than a million tokens for one survey before any trend history is considered. The economics do not work, and neither does the accuracy. Model performance degrades well before those limits are reached.
  3. The same file, two different answers: Language models are probabilistic by design. Running the same spectrum twice may result in two different diagnoses. In a research setting, that is a curiosity. When a plant is deciding whether to pull a gearbox during a scheduled outage and may later have to justify that decision in a failure investigation, it is disqualifying.
  4. Nothing to test against: Suppose a plant wanted to train a model on its own machines instead. Supervised learning requires labeled examples spectra confirmed by teardown to show inner race spalling, or a cracked gear tooth or bent shaft. Most sites have never assembled such a library, and the well-run sites have the fewest examples precisely because they prevent the failures. Building a diagnostic system that requires a fault library is building a system that cannot be validated until after it has been trusted.

The Layer That Does Work

None of this means artificial intelligence has no place in condition monitoring. It means the technology has been pointed at the wrong layer.

Consider how a competent analyst actually works. They do not stare at raw amplitude values and intuit a conclusion. They calculate. Bearing fault frequencies come from bearing geometry and shaft speed. Gear mesh frequency comes from tooth count. The analyst matches measured peaks against those calculated targets, weighs the pattern of evidence, rules out the alternatives and writes it up.

Every one of those steps except the last is deterministic physics, and deterministic physics is exactly what software has always been good at. A well-built diagnostic engine therefore has three distinct layers, and keeping them separate is the entire design.

  1. The physics layer resolves fault frequencies from machine kinematics and measured running speed, then matches spectral peaks against them. Working in orders—multiples of shaft speed—rather than in hertz is what makes this portable. A rule written for a bearing outer race defect applies to every machine in the plant, because it references the fault frequency symbolically rather than numerically.
  2. The statistics layer handles what changes over time. Robust methods matter here more than sophisticated ones. Vibration trends contain transients from process changes, poor sensor mounting and simple operator error; a conventional least-squares trend line is dragged badly off course by a single bad reading, while a median-based slope estimate ignores it entirely. This layer requires no labeled faults at all. It learns only what normal looks like for that specific asset, from that asset’s own history.
  3. The language layer is where a language model finally earns its place—and it earns it well. Given a structured set of findings from the layers below, a model can produce a clear, readable report in the plant’s working language, at the technical register the audience needs. It can answer plain-language questions across a fleet. It can help an analyst turn a hard-won diagnostic heuristic into a formal rule.
  4. The critical constraint is that the model never sees the spectrum. It receives only the conclusions and the evidence supporting them. It cannot invent a frequency it was never given.

Evidence, Not Scores

There is a second reason to prefer this structure, and in practice, it matters more than accuracy.

A diagnosis must survive contact with a maintenance planner. “Anomaly score 0.87” does not survive that meeting. What survives is “outer race fault frequency present at 109 hertz with six harmonics, no shaft-speed sidebands, envelope band energy up five-fold in 90 days, all four motor points stable.”

That second statement does more than justify itself. It tells the planner what part to order. The absence of sidebands is not a detail; a defect on a stationary outer race sits permanently in the load zone and produces harmonics without modulation, while a defect on the rotating inner race passes through the load zone once per revolution and generates shaft-speed sidebands. That single distinction separates two different repairs with different urgencies, and it is only available to a system that reasons about physics rather than pattern-matching on amplitude.

Explainability in this domain is not a regulatory nicety. It is the mechanism by which a diagnosis becomes an action.

| IMAGE 1: Vibration analysis software being used to run diagnostics (Image courtesy of BETAVIB)

Absolute Severity Still Matters

One further caution, and it applies to statistical monitoring generally: A system that only reports change will eventually mislead.

A machine whose overall velocity has risen threefold from a very low baseline may still be in excellent condition. A machine that has been running at 0.55 inches per second for two years shows no trend at all and is nonetheless well inside the zone that International Organization for Standardization (ISO) 20816 describes as sufficient to cause damage. Trend analysis answers, “What changed?” It does not answer, “Is this acceptable?” The two questions require different reference data: one from the asset’s own history, the other from published standards and the machine’s power and mounting classification.

Any monitoring approach worth adopting should answer both and should say plainly when it cannot answer either. A system that reports “insufficient data to assess this point” is more trustworthy, and ultimately more useful, than one that always produces a number.

Questions Worth Asking

For teams evaluating what is currently being marketed as AI-driven condition monitoring, the following questions separate substances from packaging:

  1. If I submit the same measurement twice, do I get the same diagnosis? If not, ask how the vendor validates anything.
  2. What was the system trained on? If the answer involves a labeled fault library, ask whether those machines resemble yours (bearing geometry, speed range and mounting all shape the signature).
  3. Show me the evidence behind one finding. A well-built system produces the frequencies, amplitudes and comparisons that led to its conclusion. A score with no derivation is not a diagnosis.
  4. What does it do when the kinematic data is missing or wrong? Bearing designations are frequently incomplete in the field. The correct behavior is to decline the diagnosis and say so.
  5. Does it distinguish “changed” from “bad”?

The Unglamorous Conclusion

The most capable condition monitoring systems being built today are not the ones that hand a spectrum to a language model. They are the ones that do the physics properly, apply careful statistics to the trends and then use a language model for the one task it is genuinely excellent at: explaining the result to a human being who has to make a decision on Monday morning.

That is a less exciting story than machines diagnosing themselves. It also happens to be the one that works.

References

  1. themanufacturer.com/articles/supply-chain-anxiety-hits-record-highs-as-geopolitical-risks-and-costs-surge
  2. weforum.org/publications/global-value-chains-outlook-2026-orchestrating-corporate-and-national-agility

Oumayma Jerbi manages the marketing team at BETAVIB, a vibration analysis and condition monitoring company headquartered in Vaudreuil-Dorion, Quebec. She works closely with BETAVIB’s engineering teams on the diagnostic platforms behind the company’s analyzers and monitoring systems. For more information, visit betavib.com.

In This Issue

Table of Contents
From the Editor
News
On the Curve
Columns
Special Section
Chemcial Processing
More Topics
Departments
Marketplace
Ad Index
Back Page
Magazine Archive

The Leading Resource for Pump Users Worldwide

Subscribe Today!

Share With Your Network: