Three tests for industrial AI before you buy it
Predictive maintenance, quality control, and energy savings are not the same use case. A NIST roadmap helps define the metric, data, and control each one must prove.
On July 25, 2026, discussion of factory AI often puts three very different tasks under one label: predicting failures, detecting defects, and reducing energy use. The promise sounds similar —greater efficiency—, but each task uses different data, changes a different decision, and needs its own metric. When they are blended, a successful vision test can be marketed as though it proved predictive maintenance. And the appetite to buy is real: according to NIST MEP, 46% of manufacturers already use AI tools such as chatbots, 55% see it as a game-changing technology, and 78% expect to increase their investment over the next two years. These are figures of intent and adoption, not of proven value — and that distinction is exactly the one this piece wants to hand back to the buyer.
The 60-second decision
Before buying a system, complete this sentence: “we want the system to recommend or perform ___ in order to improve ___, measured as ___.” If the answer is “use AI to optimize the factory,” there is not yet a testable problem.
NIST’s 2026 roadmap for smart manufacturing, published July 3, identifies these applications alongside less visible challenges: industrial data quality, integration of sensors and heterogeneous systems, and reliable, explainable operation. The issue is not simply getting a larger model. A plant needs to see what the system detects, what action it triggers, and what changes afterwards. It helps to see the whole map before choosing: NIST MEP groups factory AI into six families — machine learning (maintenance, quality, demand forecasting), robotics, computer vision (inspection and defect detection), natural language processing, predictive analytics, and digital twins. "Doing AI in the factory" can mean any of them, and none is evaluated with another's metric.
Maintenance: prediction is not repair
Predictive maintenance begins with condition signals: vibration, temperature, pressure, cycles, alarms, and maintenance records. NIST’s asset condition management framework describes the aim as assessing current condition, diagnosing it, and estimating future health so maintenance can be planned.
The measure is not model accuracy alone. It can be unplanned downtime, maintenance cost per asset, or avoided failures without creating unnecessary replacements. A correct alert that arrives too late does not prevent a breakdown; an overly sensitive one can fill a schedule with false inspections. The operator should see the signal, the threshold, and the decision: inspect, reduce load, or continue. NIST's framework orders it into three acts worth not skipping: assess the asset's current condition, diagnose why it is in that state, and estimate its future health. A vendor promising the third without demonstrating the first two sells a forecast with no foundation — and on an expensive machine, a foundationless forecast is more dangerous than none, because it invites action on it.
Quality: finding a defect is not enough
Inspection uses images, measurements, or process data to flag anomalies in a part or batch. The NIST MEP guide separates defect detection from failure prediction and resource management. Useful measures here include customer escape rate, false rejects, scrap, and inspection time.
The critical question is what happens after the result. A system can flag a weld, but someone must decide whether to repair, reject, or sample further. Keeping images and reviewed decisions exposes whether performance degrades when lighting, material, supplier, or camera changes. Without that loop, a laboratory accuracy score says little about a real line. And degradation from changing conditions is not hypothetical: a visual inspector trained on one camera and one lighting setup can collapse when the supplier changes the part's finish or someone moves a lamp. Keeping images and decisions is not bureaucracy: it is the only way to learn that the model got worse before the customer does.
Energy: optimize one variable without moving the problem
For energy, a system may forecast demand, suggest schedules, adjust setpoints, or spot anomalous consumption. A useful measure is often energy per good unit produced, not total kilowatt-hours. Stopping a line cuts consumption, but it does not make production more efficient. Quality, safety, delivery requirements, and demand peaks also matter.
NIST lists five concrete barriers: data quality and availability, high initial costs, the skills gap, privacy and cybersecurity risks, and integration with legacy systems. They are implementation conditions, not administrative details, and they explain why the 80% who say they will increase AI use is not the same as an 80% already getting results: between intent and effect stand those five doors. If an energy meter cannot be linked to product, shift, and machine state, a model may optimize a number that does not represent operating cost.
Why the plant metric is not the lab's
The thread common to all three cases — and the capability the buyer takes away — is that the figure the vendor sells is almost never the one the plant cares about. An inspection model boasts "99% accuracy"; the factory earns or loses money on something else: how many defects escape to the customer, how many good parts it over-rejects, and how much line time inspection consumes. A maintenance model boasts of catching the failure; the plant pays for unplanned downtime avoided without triggering useless replacements. Translating the vendor's metric into the operation's metric is the work, and it is where most pilots that "worked in the demo" fall apart.
That translation also has a timing trap. A failure alert is correct only if it arrives with room to act: being right that a machine will fail five minutes before it fails is, in practice, not having warned at all. And an energy metric misleads if it is not tied to product, shift, and machine state, because the easiest way to "save energy" is to produce less. The good news is that all three traps are caught with the same discipline: require the figure to be referenced to a real business unit — a part, a stoppage, a good product — and not to a laboratory percentage.
A pilot you can audit
Start with one asset, one part family, or one production cell. Record a baseline for a defined period. Specify who receives the recommendation, which actions are allowed, and when the system is reversed. Compare results with an equivalent operation and retain both successes and errors.
Do not promise a savings percentage before measuring it. The documented evidence says these methods can support maintenance, inspection, and resource management; it does not show that every plant will get the same effect. NIST MEP itself sells not a model but a method: AI-readiness assessments, training, process improvement, and implementation planning — the right order is to diagnose before buying. The durable skill is demanding the full mechanism: input data, decision, plant metric, and human review. When those four elements are clear, industrial AI stops being a label and becomes a hypothesis a factory can test.
Sources for this piece
This piece draws on 3 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.