Predictive Maintenance & Condition Monitoring: The Complete Engineering Guide
A rolling-element bearing does not fail in a second. It fails over weeks: a subsurface inclusion cracks, a spall forms on the raceway, the spall grows, debris circulates, clearance opens, temperature climbs — and somewhere in that sequence, from the very first microscopic impact, the machine was broadcasting the news as stress waves, vibration, heat, and wear particles in its oil. The failure was never sudden. The hearing was.
That is the entire premise of predictive maintenance (PdM), and it is the most economically honest branch of maintenance engineering: instead of guessing when a machine will need attention (calendar-based), or waiting for it to stop and then reacting (run-to-failure), you measure the machine's actual condition and let physics tell you how much life is left. The tools are mature, the standards are public, and the arithmetic — which this guide does end to end — is unusually favorable.
This guide covers the full condition-monitoring stack: the P-F curve that governs interval design, vibration analysis (severity, spectra, and bearing envelope diagnostics), infrared thermography, oil analysis, airborne ultrasound, and motor current signature analysis — plus the data discipline and program economics that decide whether any of it actually saves money. It pairs naturally with our guides on industrial instrumentation (the measurement chain underneath every sensor) and rolling-element bearings (the component most condition-monitoring programs are built to protect).
1. The P-F Curve: Why Machines Telegraph Their Failures
Reliability engineering formalizes the warning window as the P-F curve, introduced by the airline industry's maintenance steering groups and formalized in RCM practice:
- P — potential failure: the point where a detectable physical change appears (a defect frequency in a spectrum, a 20 °C hot spot, iron in the oil).
- F — functional failure: the point where the machine can no longer deliver its function (the pump no longer reaches head, the spindle no longer holds tolerance).
- P-F interval: the time between the two. Everything in condition monitoring exists to make this window longer and earlier.
A subtle but critical point: many failure modes do not degrade linearly. Historically, reliability research (and the original RCM studies of aircraft systems) showed that a large fraction of failure modes show no correlation between age and failure probability — they are random, not wear-out. That's the argument against pure calendar-based overhaul, and it is why condition-based tasks dominate modern maintenance programs: they monitor the actual state instead of gambling on an age distribution.
1.1 The Interval Rule You Must Respect
An on-condition (inspection) task only works if you actually inspect often enough to catch the potential-failure state before it becomes functional failure. The classic rule from RCM practice (Moubray, NAVAIR 00-25-403): the task interval must be no greater than half the P-F interval — and for safety-critical modes, stricter. The NAVAIR formalism gives the interval as:
where n is the number of inspections you manage within the P-F interval, \theta the probability that a single inspection detects the defect, and P_{acc} the acceptable probability of missing it. With a human-senses detection probability of \theta = 0.9 and an acceptable miss probability of P_{acc} = 0.001, you get n \approx 5 — one inspection every fifth of the P-F window, tighter than the folk wisdom of "half". The lesson is durable either way: if your bearing has a 6-month P-F interval and your route runs annually, you are not doing predictive maintenance — you are doing archaeology.
1.2 The Maintenance Strategy Ladder
Strategy · Trigger · Typical cost per pump hp-year, USD (Piotrowski, via NIST) · Failure predictability
Reactive (run to failure) · Failure · 18 · Zero — accept downtime
Preventive (calendar) · Time · 13 · Only for age-related modes
Predictive / condition-based · Measured condition · 9 · Weeks-months of warning
Reliability-centered (RCM + PdM + design-out) · Consequence analysis · 6 · Longest intervals, fewest failures
The same source gives a case-study ROI of 3.5:1 for moving an aerospace display system from reactive to predictive maintenance, and NIST's broader survey of US manufacturing put the perceived benefit of additional PdM adoption at USD 6.5 billion from downtime reduction alone (2016 dollars). Treat any single number with suspicion — the reporting methodology varies wildly — but the direction is consistent across every serious study: condition-based work is cheaper per operating hour than either breakdown work or calendar work.
1.3 What the Big Studies Actually Claim
Because PdM marketing is loud, it pays to know which numbers come from where — and what baseline they use:
Source · Claim · Baseline
US DOE, O&M Best Practices Guide (FEMP) · 8–12% cost savings · vs. a functioning preventive program
US DOE (same) · 30–40% savings, 70–75% fewer breakdowns, 35–45% less downtime, 20–25% more production, ROI ~10× · vs. neglected/reactive operations
McKinsey (widely cited) · 30–50% downtime reduction, 18–25% cost reduction · vs. reactive
NIST AMS 100-18 · Case studies: 15–98% maintenance-cost reduction; train wheel failure study: up to 56% cost savings · varies
The single most common error in a PdM business case is mixing baselines: "reduces downtime 50%" (McKinsey, reactive baseline) and "saves 10%" (DOE, preventive baseline) describe the same program. State your baseline or the number is noise.
1.4 The Market Context (and Why It Matters to Small Plants)
Market-size estimates for "predictive maintenance" swing widely depending on scope: roughly USD 12–19 billion in 2025–26, growing double digits, with forecasts to 2031 ranging from ~USD 24 billion (MarketsandMarkets, 11.4% CAGR) to ~USD 82 billion (Mordor, including services and platforms). Within the technology split, vibration monitoring is the largest single segment (~22.6%) and Asia-Pacific the fastest-growing region (~35% CAGR). The practical takeaway for an Indian job shop is not the headline number — it's that the sensor and software cost curve has collapsed: a battery-powered MEMS accelerometer node that cost USD 1,000 a decade ago is now a ₹10,000–25,000 commodity, and the analysis expertise needed for basic programs is a two-week certification away. The tools outgrew the excuse.
2. The Condition Monitoring Toolkit at a Glance
No single technique sees everything. A mature program layers them by the physics each one senses and correlates findings across methods:
Technique · Senses · Strongest at · Earliest warning · Typical P-F interval
Vibration (velocity, acceleration, envelope) · Dynamic force · Rotating machinery: imbalance, misalignment, looseness, bearings, gears · Bearing defects via envelope: weeks–months · Weeks to months
Thermography (IR) · Surface temperature · Electrical joints, couplings, insulation, friction/overheating · Hot connections weeks before failure · Days to months
Oil analysis · Chemistry + particles · Gearboxes, hydraulics, engines, compressors: wear mode, contamination, oil life · Wear metals before spalling · Months
Ultrasound (20–100 kHz) · High-frequency acoustic emission · Compressed-air leaks, early bearing friction, arcing/PD, steam traps · Earliest for friction & leaks · Weeks to months
Motor current signature (MCSA) · Stator current · Rotor bars, eccentricity, load problems in motors · Broken bars early in progression · Weeks to months
Performance monitoring · Process variables (pressures, flows, temps, power) · Fouling, degradation, efficiency loss of pumps/fans/heat exchangers · Gradual, monotonic · Months
The economics of layering sound like this: thermography is the highest ROI-per-rupee entry point for most plants (one camera, quarterly survey, finds both electrical and mechanical faults); vibration protects the rotating assets that produce; oil analysis is nearly free per sampled machine and covers the assets vibration struggles with (slow speed, gearboxes, hydraulics); ultrasound pays for itself on the compressed-air system alone.
3. Vibration I — The Measurement Chain and Overall Severity
Vibration is the workhorse technique and now the most standardized. Everything starts with the measurement chain — and like the instrumentation chain, it is only as honest as its weakest link.
3.1 The Accelerometer, Honestly Specified
The workhorse sensor is the IEPE piezoelectric accelerometer (Integrated Electronics Piezo-Electric — the "ICP"/"DeltaTron" family names are brand equivalents). The mundane industrial spec sheet looks like this:
- Sensitivity: 100 mV/g nominal (±10%), ~10.2 mV/(m/s²)
- Range: ±50 g typical (±80 g available)
- Frequency response: usable from ~0.5 Hz to 10 kHz (±3 dB limits typically 0.3 Hz–6 kHz on a good model)
- Mounted resonance: ~25–30 kHz — deliberately above the analysis band so the response above it rolls off predictably
- Output noise floor: ~0.2–1 mg RMS wideband
- Temperature: up to ~120 °C standard (high-temp versions to 250 °C+ for kiln/furnace work)
The detail that quietly ruins measurements is mounting. The sensor's usable frequency ceiling is set by how it is attached, because the attachment creates its own resonance:
Mounting · Practical ceiling · Notes
Hand probe / stinger · ~1 kHz · Screening only; pressure and angle vary
Magnet · ~2 kHz · Fine for most rotating-speed diagnostics
Cement / quick-connect stud · ~5–7 kHz · The practical route-survey choice
Drilled & tapped stud · ~10 kHz · Standard for diagnostics and monitoring
For a typical 1480 rpm machine, everything diagnostically interesting for imbalance and alignment sits below 500 Hz, so a magnet is fine. The moment you chase early bearing stress waves — the 500 Hz to 5 kHz band, see §5 — mounting quality becomes the measurement.
3.2 Why Velocity (mm/s RMS) Is the Severity Currency
Acceleration is the raw measurement, but machine severity is historically judged in vibration velocity, mm/s RMS, over 10–1000 Hz, because within that band the velocity produced by a given rotating-force fault varies relatively little across machine sizes and speeds, making a single scale meaningful across a plant. Modern analyzers convert between units at the touch of a button:
The integration is done in the frequency domain (divide each spectral line by 2\pi f) precisely to avoid the drift and low-frequency noise amplification that plague time-domain integration.
For slow-speed machines (below ~120 rpm — kilns, mixers, some extruders), the same physical defect produces far less velocity, and 10–1000 Hz bands miss the action; you must switch to displacement (µm peak-to-peak) and proximity probes. Different physics, different toolkit.
3.3 ISO 20816-3: The Severity Zones
The international acceptance framework is ISO 20816 (successor to the familiar ISO 10816-3). For industrial machines 15 kW–50 MW at 120–15,000 rpm, Part 3 defines evaluation zones on broadband velocity:
Zone · Verdict · Action
A · Newly commissioned condition · —
B · Acceptable for unrestricted long-term operation · —
C · Unsatisfactory for continuous operation · Plan repair; increase monitoring
D · Damage likely · Act now — reduce load or stop
The boundaries depend on machine group and support rigidity. The two most common rows for plant equipment (rigid support):
Machine group · A/B · B/C · C/D
Group 1 — large (>300 kW or shaft height >315 mm) · 2.3 mm/s · 4.5 mm/s · 7.1 mm/s
Group 2 — medium (15–300 kW) · 1.4 mm/s · 2.8 mm/s · 4.5 mm/s
Flexibly-mounted machines get boundary values ~25% higher (foundation flexibility isolates, but also amplifies at resonance — the standard treats them separately). For journal-bearing machines, the accompanying shaft-vibration annex works in displacement, with zone boundaries scaling as 4800/\sqrt{n} µm (A/B) through 13200/\sqrt{n} µm (C/D), where n is rpm — at 1500 rpm, an A/B boundary of 124 µm peak-to-peak.
3.4 Worked Example — Judging a Pump Motor
A 22 kW process pump motor, rigidly mounted, reads 3.6 mm/s RMS at the drive-end bearing housing. Group 2 boundaries are 1.4 / 2.8 / 4.5 mm/s. The reading sits in Zone C — "unsatisfactory for continuous operation": schedule diagnosis, don't wait for the next shutdown. The same reading on a large Group 1 machine would be comfortably Zone B. This is why "what's a bad vibration level?" has no universal answer — the standard is a three-variable function (power class, support, speed regime), and the machine's own baseline trend is a fourth.
Two field rules that keep the zones honest:
- Check zone validity against OEM limits. API/ISO machinery-specific standards (e.g., API 670 for turbomachinery protection) override the generic framework.
- Never alarm on a single absolute number. A machine that has run 12 years at 1.1 mm/s and reads 2.9 mm/s after a rebuild has a problem, even though 2.9 is "only" Zone B for its class. Trend relative to its own baseline first; zones are a sanity frame, not a trip wire.
4. Vibration II — The Spectrum Is the Fingerprint
Overall severity tells you whether something is wrong. The frequency spectrum tells you what. Once you internalize a dozen signatures, a spectrum reads like a medical report.
4.1 FFT Arithmetic You Must Get Right
The FFT converts a time record into a spectrum; four parameters define the measurement:
- F_{max} — set it above the highest fault frequency you care about (2.5× is the common floor; for bearing work, 10–20× shaft speed is typical).
- Nyquist: sampling must exceed 2 F_{max}; analyzers oversample (~2.56×) and filter, because a hot signal above F_{max} folds back and aliases into a false peak. Aliased peaks are the classic newbie ghost.
- Resolution \Delta f — match it to the smallest frequency separation you must resolve. For a 1480 rpm machine (f_r = 24.7 Hz), 400 lines at 1 kHz Fmax (2.5 Hz bins) is fine for imbalance and gear mesh; 3200 lines (0.31 Hz) is the diagnostic default; slow-speed machines and MCSA slip sidebands demand 0.01–0.1 Hz — 12,800 lines and longer records.
- Averaging and windows — 4–8 linear averages to suppress random noise; a Hanning window to stop spectral leakage from non-bin-centered peaks. Long records (tens of seconds) have a hidden cost: speed drift smears peaks, so for fine resolution use tachometer-triggered or order-tracked measurements.
4.2 The Classic Fault Signature Table
This table is the working core of vibration diagnostics — amplitudes are read in velocity, except where noted. The columns encode the two things that matter: which frequencies appear, and in which direction (radial vs axial):
Fault · 1× · 2× · Other signature · Direction · Notes
Imbalance · Dominant · Small · — · Radial · Amplitude ∝ rpm²; phase stable
Misalignment (parallel) · Present · Dominant · Sometimes 1× axial · Radial + axial · 2× axial is the tell
Angular misalignment · Present · High · — · Strongly axial · Compare H/V/Ax amplitudes
Mechanical looseness · High · High · 3×–10× "harmonic forest", raised noise floor · Radial · Sub-harmonics (0.5×) when severe
Soft foot · 1×, 2× · Moderate · Axial at foot · Varies · Confirm by loosening one bolt at a time
Belt drive · Present · Strong · Belt pass frequency f_B = \frac{\pi D n}{L} · Radial · Worn belts: 3–4× of f_B
Gear mesh · — · — · GMF = z·f_shaft + harmonics · Radial/axial · Sidebands ±f_shaft = tooth defect
Bearing defects · — · — · Non-integer defect frequencies (§5) · Radial · Usually invisible in early stage
Electrical (stator) · — · Dominant · 2× line frequency (100 Hz @ 50 Hz) · Radial + axial · Collapses when power removed
Rotor bar / air-gap · 1× · — · Sidebands at ±slip frequency around 1×; 2×fl sidebands · Radial · Cross-check with MCSA (§9)
Flow/aero · — · — · Blade/vane pass (blades × rpm) · Radial/axial · Varies with flow, not speed
The single most valuable habit in spectrum reading: watch the integer multiples of running speed first (1×, 2×, 3× — mechanical forcing), then look for non-integer families (bearing defect orders, gear mesh), then check line-frequency relatives (100 Hz, 200 Hz, ±slip sidebands). The mental checklist takes ninety seconds and resolves the large majority of cases.
4.3 Gearbox Case — Mesh Frequency and Its Sidebands
A 24-tooth pinion at 1470 rpm has shaft frequency 24.5 Hz and gear mesh frequency:
A healthy gearbox shows GMF plus low, orderly harmonics. Then the diagnostic branch:
- GMF amplitude rising with sidebands spaced at 1× shaft speed (±24.5 Hz) — amplitude modulation at shaft rate = a localized tooth problem (crack, spall, single damaged tooth) on that gear.
- Whole GMF family rising uniformly with low sidebands — distributed wear or misalignment of the mesh rather than one tooth.
- Hunting tooth frequency — for gear pairs, f_{HT} = GMF / \mathrm{LCM}(z_1, z_2) describes how often a given pair of teeth re-meets; damage concentrated at that rate points to a specific pair. (Our guide to gears and gearboxes covers the bearing-level view of transmission design.)
- Ghost frequencies from manufacturing (indexing errors) appear at non-mesh orders and never grow — flag them once, then ignore.
4.4 Diagnostic Discipline
Three rules separate diagnosis from interpretation:
- Confirm with operating condition changes. Electrical frequencies (100 Hz etc.) vanish when the motor is de-energized; flow-induced peaks change with valve position, not speed; mechanical forcing tracks rpm. Each experiment halves the hypothesis space.
- Check all three directions. A purely radial 1× is imbalance; the same 1× with heavy axial could be angular misalignment or a bent shaft. Sensors go Radial-Horizontal, Radial-Vertical, and Axial on every bearing — no exceptions on critical assets.
- Trust ratios, not single peaks. Compare 2×:1× (alignment vs balance), compare defect families across bearings (which one leads?), compare current vs last survey (trend slope, not level).
5. Vibration III — Bearing Diagnostics: Defect Frequencies and Envelope Analysis
Roughly half of rotating-machine failures trace to bearings and lubrication — the single highest-value target in any PdM program. Their vibration physics has two layers: geometry (which frequencies appear) and energy (why you can't see them early in a normal spectrum).
5.1 The Defect Frequency Formulas
For a rolling-element bearing with n rolling elements, element diameter d, pitch diameter D, contact angle \theta, and shaft frequency f_r (Hz), the four kinematic defect frequencies are:
where BPFO/BPFI are the ball-pass frequencies for outer/inner race defects, BSF is the ball/roller spin frequency, and FTF is the fundamental train (cage) frequency. Two free checks catch most arithmetic errors:
- \text{BPFO} + \text{BPFI} = n\, f_r (exact, for zero slip).
- \text{BPFO} = n \cdot \text{FTF} (the outer race sees every element once per cage revolution).
And two useful approximations when the bearing drawing isn't at hand: for most ball bearings d/D \approx 0.2, so \text{BPFO} \approx 0.4\, n f_r and \text{BPFI} \approx 0.6\, n f_r. Use them to place cursors, never to write reports.
5.2 Worked Example — a 6205-Class Bearing at 1480 rpm
Take a ubiquitous 6205 deep-groove bearing: 9 balls, ball diameter 7.94 mm, pitch diameter 39.04 mm, contact angle 0°. At 1480 rpm, f_r = 24.67 Hz and d/D = 0.2034:
Quantity · Frequency · Orders of running speed
BPFO (outer race) · 88.4 Hz · 3.58×
BPFI (inner race) · 133.6 Hz · 5.42×
BSF (ball) · 58.1 Hz · 2.36×
FTF (cage) · 9.8 Hz · 0.398×
Checks: 88.4 + 133.6 = 222.0 = 9 \times 24.67 ✓. Note the orders are non-integer — 3.58×, 5.42×, 2.36×, 0.398× — and that is diagnostic gold: it's how a defect family is distinguished from running-speed harmonics (1×, 2×, 3×…) in a spectrum. Every real bearing has its own card of these numbers; the manufacturer publishes factors, and every analyzer has them in a library. For real diagnoses, use the manufacturer's values — slippage shifts the true rate 1–2% from kinematics, harmless for identification, fatal for sloppy "exact frequency" claims.
5.3 Why Early Bearing Damage Hides — and How Envelope Analysis Finds It
At the earliest stage, a spall is microscopic. Each impact is tiny, lasts microseconds, and is broadband — its energy is spread across kilohertz, not concentrated at the elegant 88.4 Hz above. Worse, discrete impacts excite the structural resonances of the bearing housing (typically 500 Hz–5 kHz, even higher on stiff assemblies), so the energy lands in a diffuse high-frequency zone where classic spectrum analysis sees only a mildly raised noise floor. RMS velocity may be utterly normal — the machine sails through its ISO zone while a spall grows.
Envelope analysis (demodulation, also called the bearing high-frequency technique) is the standard extraction method:
- Band-pass the raw acceleration signal around a structural resonance band — say 1 kHz–5 kHz — where the impacts ring strongest.
- Rectify and demodulate: extract the signal's envelope (Hilbert transform or peak detection), discarding the high-frequency carrier.
- FFT the envelope. The low-frequency "beat" of repeated impacts re-emerges — and now the defect frequencies (88.4, 133.6, 58.1 Hz and their spacing) stand clear of the noise.
The practical result: envelope spectra can show the defect family weeks to months before it appears in a velocity spectrum. Amplitude is typically reported in acceleration units (gE — "g's envelope") and interpreted by trend, not threshold: the difference between 0.2 gE steady and 0.2 → 0.9 gE in three surveys is the alarm.
The sideband logic (from §4's chessboard, specialized to bearings):
- BPFI family with sidebands spaced at 1× shaft speed — inner-race defect. The defect rotates with the shaft, so it passes in and out of the load zone once per revolution, amplitude-modulating the impacts.
- BPFO family with no 1× sidebands — outer-race defect. The outer race is stationary relative to the load zone; the impacts stay uniform.
- BSF with FTF sidebands — rolling-element defect (2×BSF often dominant).
5.4 The Earliest Detectors: Shock Pulse and Stress Waves
Two specialist techniques see even earlier than envelope analysis, both rooted in the same physics — the stress wave emitted by each impact travels away from the bearing at ultrasonic frequencies:
- Shock pulse method (SPM): a transducer deliberately resonating around ~32 kHz converts each impact into a decaying burst; the instrument reports dBm (maximum impact level) and dBc (background carpet level). A healthy, well-lubricated bearing shows low dBm with dBc close behind; damage widens the dBm−dBc spread and drives dBm upward. Bonus: the reading is far less sensitive to machine speed than classic vibration.
- Acoustic/stress-wave gE sensors: high-frequency accelerometers with band-limited analytics for the same purpose, integrated in modern route analyzers.
Practical sequencing: ultrasound or SPM (earliest, lubrication-sensitive) → envelope (defect frequencies emerge) → velocity spectrum (defect families, sidebands visible) → overall level (late stage — broadband energy, harmonics, 1× growth, audible noise). By the time overall velocity is climbing, you're no longer scheduling work — you're managing an incident.
6. Thermography — Seeing Heat Before Failure
Every surface above absolute zero radiates infrared; a thermal camera images the 7.5–14 µm long-wave band and converts radiance into a temperature map, governed by the Stefan-Boltzmann law:
The entire technique lives or dies on emissivity \varepsilon — how efficiently a surface radiates compared to a perfect blackbody. Get it wrong and the math is quietly wrong downstream:
Surface · Emissivity \varepsilon · Trap
Polished copper · ~0.03 · Reads absurdly cold — it's a mirror in LWIR
Oxidized copper · ~0.78 · The dull version of the same busbar is measurable
Painted / oxidized steel · ~0.95 · What you actually want to point the camera at
Shiny aluminum, galvanized · 0.05–0.25 · Unreliable unless taped
Misestimating \varepsilon by ±0.1 produces temperature errors up to ~15% (per ISO 18434-1 guidance) — so field practice is: put a strip of high-emissivity tape or matte paint on critical targets, and read relative differences (phase-to-phase, unit-to-unit) rather than trusting absolute numbers on shiny metal.
6.1 Standards and the ΔT Decision Table
The framework standards are ISO 18434-1 (machine condition monitoring and diagnostics using infrared thermography) with personnel certification under ISO 18436-7. The commercially most consequential development: NFPA 70B (2023 edition) converted electrical maintenance from "recommended practice" into standard language — infrared inspection of electrical equipment is now a documented, load-dependent, repeatable program requirement in compliant facilities, not an occasional walk-around.
The widely cited NETA suggested-action criteria, adapted:
ΔT (component vs. reference) · Interpretation · Action
1–3 °C · Possible deficiency · Investigate; compare to trend
4–15 °C · Probable deficiency · Repair as time permits
16–40 °C · Serious · Schedule repair promptly; increase survey frequency
> 40 °C · Critical · Repair immediately
Two cautions keep the table honest. First, context dominates magnitude: a 10 °C rise on a 100 A fuse holder is far more consequential than the same rise on a large busbar — the fuse's smaller mass sheds less heat and runs closer to its failure envelope. Second, the physics is I²R: heat scales with the square of current, so a defect only shows itself under load — survey at ≥40% of rated load or you are photographing a healthy-looking lie.
6.2 Survey Discipline
- Load first: run equipment at normal (or ≥40%) load; note it in the report.
- Emissivity and reflection: set \varepsilon per target; watch for reflected heat from adjacent hot surfaces or the sun (outdoor work: no direct sun, wind < ~15 km/h).
- Spot size: know your IFOV; you need the target to cover at least 3×3 pixels for a trustworthy reading — a 320×240 detector with a typical 25° lens at 2 m projects each pixel onto ~2.8 mm of target, so the smallest reliable spot is ~8–9 mm, which is exactly the scale that matters for terminal-level work.
- Consistency: same route, same angles, same load points, saved reports — thermography is a trend tool as much as a find-the-hot-spot tool.
- Safety: energized-panel work means PPE per arc-flash boundaries, always.
6.3 Applications and Economics
- Electrical (the killer app): loose terminations, corroded lugs, overloaded breakers, failing CTs, unbalanced phases. A loose connection is just a resistor: every extra milliohm becomes watts, then degrees, then carbon, then an arc flash. Finding one during a quarterly survey costs a re-termination; missing it costs a shutdown and possibly an incident.
- Mechanical: bearing overtemperature (though vibration and ultrasound see it earlier), coupling misalignment via differential coupling-leg heat, belt slip, steam trap failures, furnace refractory hot spots, tank levels.
- Camera economics (indicative India, 2026): 160×120 entry cameras ₹20,000–60,000 handle panel surveys; 320×240 class instruments with ~0.08 K thermal sensitivity run ₹1.2–4 lakh and hold calibration well enough for report-grade work. One quarterly survey of a 120-panel plant is a two-day job that routinely pays for the instrument within a year.
7. Oil Analysis — Reading Wear in the Lubricant
Oil is the machine's bloodstream: it carries both its own chemistry (is the lubricant still fit for service?) and the machine's debris (is anything wearing?). A structured program answers three separate questions with three separate test families:
Question · Key tests · Methods (standards)
Is the oil still fit? · Viscosity @ 40 °C, TAN/TBN, oxidation & nitration · ASTM D445, D664/D2896, FTIR (E2412, D7412, D7624)
Is the machine wearing? · Wear metals, PQ index, particle morphology · ICP-OES (ASTM D5185), PQ (D8184), analytical ferrography
Is the oil contaminated? · Particle counts, water, fuel/glycol · ISO 4406, Karl Fischer (D6304), GC screens
7.1 ISO 4406 — The Cleanliness Code Decoded
ISO 4406 reports three numbers — particle counts at ≥4, ≥6, and ≥14 µm(c) per millilitre, each mapped to a logarithmic code. The scale doubles every step:
Code · Particles/mL (≥4 µm) · Code · Particles/mL (≥6 µm) · Code · Particles/mL (≥14 µm)
22 · 20,000–40,000 · 20 · 5,000–10,000 · 16 · 320–640
20 · 5,000–10,000 · 18 · 1,300–2,500 · 14 · 80–160
18 · 1,300–2,500 · 16 · 320–640 · 13 · 40–80
16 · 320–640 · 14 · 80–160 · 11 · 10–20
So a report of 18/16/13 means 1,300–2,500 particles ≥4 µm, 320–640 ≥6 µm, and 40–80 ≥14 µm per mL. Typical targets: hydraulic systems below ~140 bar: 18/16/13; high-performance servovalve systems: 16/14/11 or tighter. Fresh oil from the drum is often worse than the target — never assume "new" means clean; filter on fill or accept early filter loading. And note the trap of the code scale: one step is a factor of two, so a shift from 18/16/13 to 21/19/16 is not "a bit dirtier" — it is ~8× the contamination, i.e., an active ingress path (breather, seal, or fill point).
7.2 Wear Metal Decoding
ICP-OES gives parts-per-million concentrations of ~20 elements; the art is reading them as a story:
Element · Common sources · What a rise suggests
Fe · Gears, shafts, liners, hydraulic cylinders, rust · General steel wear; least specific, most trended
Cu · Bushings, bearing cages, oil coolers, thrust washers · Brass/bronze wear or cooler corrosion
Cr · Piston rings, hard-chromed rods, some bearings · Ring/plating wear — often abrasive or corrosive attack
Al · Pistons, bearing cages, pump bodies · Light-alloy wear; on engines also dust/dirt
Si · Dust ingress · Abrasive contamination (check breathers and seals)
Na, K · Coolant, process water · Coolant ingress (engines) or steam contamination (turbines)
Pb, Sn · Bearing overlays, solders · White-metal bearing wear — treat seriously
Three rules make elemental data useful rather than decorative:
- Trend rates, not single readings. Absolute thresholds differ per machine; the alarm is "Fe doubled in one interval," not "Fe > 100 ppm."
- Cross-check with PQ index and ferrography. ICP is blind above roughly 5–8 µm — the big, chunky failure particles slip past it, which is exactly why large failures can show "normal" ICP iron. PQ responds to bulk ferrous debris; ferrography shows particle shape: cutting wear (machining-like chips), fatigue (spherical/spall flakes), sliding wear (fine platelets).
- The sample is half the analysis. Live-zone (flowing, before the filter, never from a dead drain port), machine at operating temperature, same location and method every time, recorded hours and top-up volumes. A perfect laboratory on a bad sample measures nothing.
7.3 Intervals and Costs
Start critical gearboxes and hydraulic systems at 500–1,000 operating hours, then stretch or shorten by observed stability; engines and compressors follow OEM schedules. Indicative Indian lab pricing (2026): a standard kit — viscosity, water, ISO 4406, ICP elemental suite — runs ₹800–2,000 per sample; an advanced panel adding FTIR, PQ, and ferrography screening runs ₹2,500–6,000. Against the cost of a scuffed ₹8 lakh gearbox, the arithmetic is not close. (Our hydraulics guide covers the system-side view of fluid cleanliness targets.)
8. Ultrasound — Hearing Leaks, Friction, and Arcing
Airborne ultrasound instruments listen in the 20–100 kHz band — above audible machine noise but rich in the emissions of turbulence, friction, and electrical discharge. Turbulent flow through a small orifice (a leak) generates broadband ultrasound; the instrument heterodynes it down to an audible hiss whose loudness scales with proximity. Instruments report levels in dBµV (SDT reference: 0 dB = 1 µV; a typical leak checker spans −6 to 99.9 dBµV with 35–42 kHz measurement bandwidth). In noisy steel plants, the classic narrowband 40 kHz listen can be drowned by furnace and drive noise — broadband instruments (20–100 kHz with spectrogram display) or shifting the listening carrier to ~70 kHz restores the signal-to-noise ratio, a trick worth knowing before declaring a system "too noisy to survey."
8.1 The Compressed-Air Case — Pure ROI Math
Compressed air is the most expensive utility in a shop (thermal efficiency of compression lands in single digits to low teens), which makes leaks a silent tax. The arithmetic for a single 1 mm hole at 6 bar:
with \dot{V} \approx 1\ \text{L/s} = 3.6\ \text{m}^3/\text{h} (the standard leak table rate for 1 mm at 6 bar), specific energy e \approx 0.11\ \text{kWh/m}^3, and electricity at ₹8.5/kWh:
That is ₹29,500 per year from one millimetre hole at 6 bar. A 3 mm hole passes roughly 10× the flow: ~₹2.9 lakh per year from a defect you can barely see. Typical plants leak 20–30% of production; a 75 kW compressor running at ~70% average load consumes about 460,000 kWh/yr (₹39 lakh); the leak share of that is ₹8–12 lakh per year — money literally vented to atmosphere. An ultrasound detector (₹1.5–5 lakh) that finds and helps fix leaks in a couple of night-time surveys pays for itself within months, every year, forever. Add the other ultrasound duties — bearing lubrication audits (add grease until the dB reading bottoms out, stop before churning raises it again), slow-speed bearings that vibration can't see, steam trap condition, valve pass-by, and electrical arcing/partial-discharge checks in panels — and no plant tool has a better ratio of findings to money. (The compressed-air guide covers system-level leak economics and piping design.)
9. Motor Current Signature Analysis and Electrical Diagnostics
A healthy induction motor is, electrically speaking, a beautifully predictable machine — so when its magnetic symmetry degrades, the stator current carries the evidence. Motor current signature analysis (MCSA) clamps a current transformer around one phase, digitizes the waveform, and FFTs it with punishing resolution. The marquee application — broken rotor bars — produces characteristic sidebands around the supply frequency:
For a 4-pole motor on 50 Hz running at 1470 rpm, slip s = (1500-1470)/1500 = 0.02: sidebands at 48 and 52 Hz — only 2 Hz from the fundamental. Resolving them requires 0.01–0.05 Hz frequency bins (records of tens of seconds), a steady load of ~70–80% or better, and patience: at light load, slip shrinks and the sidebands hide inside the fundamental's skirt. Severity is judged by sideband amplitude relative to the fundamental — in a documented utility case, a 2.4 MW motor was pulled on the strength of −32 to −37 dB sidebands and rotor damage was confirmed on inspection.
What else the current spectrum carries:
- Air-gap eccentricity: families tied to rotor mechanical speed around the supply and slot harmonics — static vs dynamic eccentricity have distinguishable patterns.
- Stator winding asymmetry: sidebands that survive load changes (unlike load-induced effects).
- Load problems: oscillating loads (reciprocating compressors, torn belts) modulate current in ways a power logger sees cheaply.
Practical economics: the incremental cost of MCSA is close to zero on already-instrumented machines — modern protection relays and VFDs measure current continuously. For motors above ~100 kW, a waveform-capture-capable relay plus analysis software is the cheapest new diagnostics channel a plant can add. Caveats: inverter-fed motors complicate the spectrum (carrier-frequency artifacts, and the fault signatures change), so capture on the motor side of the drive and compare against a healthy baseline of the same machine. And cross-check everything against vibration — a bearing outer-race defect on the drive end shows in both spectra, which is exactly the confirmation confidence you want before pulling a 200 kW motor. (See our electric motors guide for the machine-side context.)
For the broader electrical estate, the offline toolkit rounds out the picture: insulation resistance and polarization index (IEEE 43) on critical motor windings, surge/partial-discharge tests on MV assets, and contact-resistance checks on breakers — paired with the thermography of §6, this covers the failure modes that account for the majority of unplanned electrical outages.
10. Turning Signals into Decisions — Baselines, Alarms, and the Analytics Layer
Data is cheap; decisions are the product. The discipline that separates a working program from a sensor installation:
Baseline or bust. Capture "known good" data at commissioning and after every rebuild — machine hot, loaded, at normal operating point. Without it, you are guessing what "bad" looks like: a large motor's brand-new 2.6 mm/s is somebody else's emergency.
Alarm philosophy — three layers, in order:
- Trend vs. baseline (the primary trigger): e.g., alert when velocity RMS reaches 2× the stable baseline, or when any band shows a step change > 50% sustained across two surveys.
- Rate of change: doubling within weeks matters even if absolute values are "acceptable" — a bearing spall's early envelope growth is often exponential, not linear.
- Absolute frames (ISO 20816 zones, OEM limits): the sanity boundary that keeps the program defensible and comparable across machines.
Define the action for each alarm level before creating the alarm. An alarm with no defined response is just anxiety, and it will be ignored within a quarter.
The condition indicators and their traps:
Indicator · Reads · Interpretation notes
Velocity RMS (mm/s) · Overall severity · Use ISO zones + own baseline; late for bearings
Peak acceleration (g) · Impact energy · Sensitive early, noisy late
Crest factor (peak/RMS) · Waveform spikiness · Sine = 1.414; healthy bearing ~2–3; rises early, then falls as late-stage signal densifies — never read alone
Kurtosis · Statistical peakedness · Gaussian noise = 3; spiky defects push it > 4–5, then it normalizes in the endgame
Envelope RMS (gE) · Bearing-specific energy · The workhorse early indicator; trend it
Temperature · Friction/overload · Slow-moving but brutally definitive
The crest factor/kurtosis "rise then fall" behavior is the classic trap of statistical indicators: they detect damage onset brilliantly and damage severity poorly. Envelope trending carries the diagnosis; statistics raise the flag.
The analytics layer, without the hype. Machine learning earns its keep in three jobs: anomaly detection on stable-duty machines, automated screening of route data (flagging the 5 machines that changed among 500), and remaining-useful-life estimates where failure data exists. What it cannot do yet, for most plants: conjure diagnosis from a per-machine data trickle (ML models starve without failure examples), or replace the physics reasoning in §4–§5. The pragmatic architecture: physics-based rules are authoritative; ML is a second opinion and a screening tool; a human expert reviews monthly. Start capturing raw waveforms now — storage is cheap; data you didn't save can't train anything later.
Wireless vs. route. Battery MEMS nodes (₹8–25k/node, 3–5 year battery at hourly capture) transformed coverage: hard-to-reach, hazardous-area, and dozens-of-small-motors cases that never justified a manual route. But check the sensor's bandwidth and dynamic range against your targets — a budget MEMS node covering ~10 Hz–1 kHz cannot see the 2–5 kHz bearing resonances, which is precisely what envelope analysis needs. The mature split: wireless for coverage and alerting, portable route instruments for deep diagnosis, with routes on the critical machinery where expert eyes earn their keep.
11. The Economics — What Uptime Actually Costs in India
11.1 A Worked Downtime Event
A machining shop's horizontal machining center loses a spindle bearing. The machine produces at a contribution margin of ₹2,500 per worked hour, 16 productive hours per day. The event unfolds as: 4 days of diagnosis and gentle running while convincing everyone it's serious, 5 days waiting with the machine stopped plus 2 days of rebuild once the spindle returns — call it 7 lost production days (₹2.8 lakh of margin gone), a ₹1.2 lakh rebuild and freight, and the scrap produced in the run-up. Round the event to ₹4.3 lakh, plus the second-order damage (missed delivery, expedited freight on other jobs, the customer call nobody wants to make).
The counterfactual: a monthly vibration route reading costing ₹3,000 per visit, or a ₹20,000 wireless node watching the spindle's envelope trend, would have flagged the defect family 8–12 weeks out. Converted into a planned repair at the next scheduled idleness, the same job costs a fraction of that and produces zero lost days. One avoided event pays for years of monitoring on that machine — this is why the DOE's aggregate claim of 8–12% savings over a preventive program (and 30–40% over reactive operations) survives scrutiny: the individual arithmetic is mundane.
11.2 Indicative Program Costs (India, 2026)
Item · Indicative range
Handheld vibration meter (screening) · ₹15,000–40,000
Portable vibration analyzer + software (route-grade, Cat II class) · ₹1.5–8 lakh
Wireless accelerometer node · ₹8,000–25,000
Wireless gateway · ₹40,000–1.5 lakh
Cloud/analytics platform · ₹20,000–60,000/year
Thermal camera, 160×120 · ₹20,000–60,000
Thermal camera, 320×240, ~0.08 K · ₹1.2–4 lakh
Ultrasound detector · ₹1.5–5 lakh
Oil sample (standard kit / advanced panel) · ₹800–2,000 / ₹2,500–6,000
ISO 18436 Cat I training · ₹40,000–80,000
Baseline survey by external consultant · ₹5,000–15,000 per machine
11.3 Program ROI, Worked
A 30-machine critical subset, year one: 20 wireless nodes (₹3 lakh) + gateway (₹1 lakh) + one route analyzer (₹2.5 lakh) + Cat I training (₹60k) + software (₹40k) ≈ ₹7.5 lakh. Break-even: under two avoided events of the §11.1 class. Steady state adds the second-order savings that never make the spreadsheet cleanly — the Piotrowski pump data (via NIST) puts the difference between reactive and predictive maintenance at roughly USD 9 per hp-year, which for a 22 kW (30 hp) pump is about ₹24,000 per year per pump; a plant with 40 such pumps is looking at ~₹95 lakh/yr of theoretical differences distributed across the maintenance accounts. The honest caveats: these models assume you actually act on findings, and programs die from organizational failures — alarm noise, unclosed findings, no baseline — long before they die from technical limits. The sensors are the easy part.
12. Implementation Roadmap for a 10–200 Machine Plant
- Criticality pass first. Rank assets by downtime cost per hour, safety impact, and spare lead time. In a typical job shop, 10–20 machines carry most of the consequential risk; they get the program, everything else stays on preventive maintenance plus operator senses.
- Baselines before alarms. Capture good, hot, loaded data at commissioning and after every rebuild. Thirty days of steady-state readings beats any generic alarm value.
- Start one technique, done properly. A monthly vibration route on the top 10 machines + a quarterly thermal survey of panels + 2–4 oil samples/year on gearboxes and hydraulic systems is a legitimate program. It will find things in the first quarter.
- Route discipline. Same points (paint them), same directions, same load conditions, logged in the CMMS. Data discipline is the moat.
- Design alarms with actions. Baseline delta + ISO frame + rate of change; every level maps to a named response.
- Train one person to ISO 18436-2 Category I and keep a contract analyst for escalation. Cat I finds the routine signal; complex diagnoses (gearboxes, slow-speed) get professional eyes quarterly or on demand.
- Extend by economics, not fashion. Wireless nodes to inaccessible or hazardous assets; MCSA where motors exceed ~100 kW; ultrasound where compressed air bills hurt.
- Fix findings fast. A program that doesn't repair what it finds loses credibility faster than it loses bearings.
- Review quarterly. Findings, fixes, rupees saved, and — critically — alarm tuning to kill false positives. Nobody trusts a system that cries wolf.
Pitfalls that kill programs: over-sensoring everything on day one; skipping lubrication fundamentals (most "bearing failures" are lubrication failures — fix grease practice before buying dashboards); using velocity-only monitoring on slow-speed machines; running trends against shifted measurement points (the classic false alarm); and treating a software subscription as the program.
13. Where Condition Monitoring Meets Fabrication
Every monitoring program generates a quiet stream of fabrication demand that most plants discover halfway through rollout: machined sensor mounting pads and bosses (stud-mounting needs a flat, tapped, precisely faced surface — often impossible on an existing casting without a custom adapter), stainless brackets for wireless nodes in wash-down or outdoor areas, enclosures and sub-panels for gateways and junction boxes, cable management hardware, calibration fixtures, and the repair pipeline that PdM triggers — refurbished shafts, machined bearing housings, custom sleeves and spacers when lead times on OEM spares run to weeks. None of it is glamorous; all of it is precision work with drawings and tolerances attached, which is precisely the class of low-volume fabrication that marketplaces like FabFlow connect to vetted shops every day. If a reliability rollout leaves you with a bill of brackets, adapters, and housings — that's a solved problem.
The Reliability Engineer's Checklist
- Baseline every critical machine hot and loaded — and re-baseline after every rebuild.
- Inspect at half the P-F interval or less; for safety-related modes, use the probability math, not folklore.
- Same points, same directions, same load on every vibration collection — RF and Axial on every bearing.
- Velocity for severity (ISO 20816), envelope for bearings, displacement for slow speed. One scale does not fit all machines.
- Keep a defect-frequency library for your top 50 bearings — manufacturer factors, not approximations.
- Thermal surveys at ≥40% load, high-ε targets, quarterly on electrical panels.
- Oil: live-zone samples, trend rates not absolutes, ISO 4406 targets set by system pressure class.
- Night-idle leak surveys quarterly; fix by rupees, not by discovery order.
- Never create an alarm without a defined action; tune quarterly.
- Close the loop — every finding gets a work order or a documented engineering decision to accept it.
Machines are honest witnesses; they testify continuously, in vibration, heat, chemistry, and sound. Condition monitoring is simply the discipline of showing up to hear the testimony early enough to do something about it — with arithmetic that works.
Standards and sources referenced: ISO 17359 and ISO 13373-1 (condition monitoring, vibration); ISO 20816-3:2022 (successor to ISO 10816-3; vibration evaluation zones); ISO 18436-2 / -7 (vibration analyst and thermographer certification); ISO 18434-1:2018 (infrared thermography for machine condition); ISO 29821-1 (airborne ultrasound); ISO 4406:2021 (oil cleanliness); NFPA 70B (2023) and NETA ATS (electrical maintenance and thermography criteria); ASTM D445, D5185, D8184, D6304, D7412, D7624, E2412 (lubricant testing); IEEE 43 (insulation testing); DoDM 4151.22 (RCM, P-F interval); US DOE FEMP O&M Best Practices Guide (savings baselines); NIST AMS 100-18 (maintenance economics); vendor application notes from SDT/SONOTEC (ultrasound leak detection). Market figures per Mordor Intelligence, MarketsandMarkets, and Astute Analytica (estimates vary by scope). Prices are indicative Indian street levels as of 2026 and exclude GST.