Article

You Are the Evaluation Body: Testing AI Against a Physical Process

There is no outside body evaluating AI against your production line, and the national testbed that just launched is not pointed at manufacturing. So the acceptance test has to be one you build: a frozen baseline, a read-only run across a full cycle, and a number written down before you look.

Nobody is coming to benchmark your production line for you. There is no outside body sitting between a model and your press, your oven, your coating line, and no independent evaluation a vendor has to clear before handing you something that moves a physical asset. That job is unclaimed. If you don't do it, it doesn't get done.

The map of who is responsible

Serious evaluation work is happening. On July 27, 2026, NIST announced its Artificial Intelligence Technology Evaluation (AITE) program — a sequestered testbed for evaluating AI model performance, initially focused on image analysis with vision-language models across quantum science, genomics, and public safety imaging. Sequestered data, blind tasks, real methodological discipline. It is the kind of program I want to exist.

Now read that list as somebody who runs a plant. A stamping press is not on it. Neither is a curing oven, or a coating line running a material that behaves differently in August than in February.

And this isn't a queueing problem where manufacturing is simply next. NIST's own AI for Manufacturing workshop lists “the development and validation of AI-driven operations, maintenance, or control strategies” and “the measurement and evaluation of the impact of AI on physical processes and assets” among its topics of interest. Those are framed as open questions. Not settled methods, not a standard you can hand your quality manager — open measurement-science questions, from the institution whose entire job is measurement science.

That is not a knock on NIST. It is a map of who is responsible, and the answer for your process is you.

Two different objects

There is a general-purpose evaluation: can this model do this class of task at all, scored against a held-out set somebody else built. It is portable, and it is useful for choosing what to trial.

Then there is your process acceptance test: will this specific system, on this line, with your material, on your worst shift, make parts you can ship. It exists only if you build it, and it is the only one that decides whether the line runs. Confusing them is how a leaderboard result ends up doing work it was never designed to do.

Make them finish the sentence

The word to stop accepting at face value is validated. When a vendor says a system is validated, the sentence is missing its second half. Validated against what — their cell, their material, their assumed variance, which in practice is a distribution somebody typed into a configuration file.

Nobody outside your building has your changeovers, your second shift in July, or the stretch where a supplier quietly changed a coating and scrap moved before anyone noticed. Ask out loud, in the room, and make them finish the sentence. If the honest answer is “validated in simulation, under our assumptions,” that's fine — it's still useful. It just isn't your acceptance test.

The rig, in three unglamorous pieces

Freeze a baseline, before the vendor is in the building — one captured after the project starts is already contaminated by everyone's motivation for it to work. Capture cycle time, scrap, rework, unplanned downtime, changeover time — and the human hours nobody logs, the ones spent babysitting the process when it drifts. That category is usually where the real gain is, or isn't. Practically: a dated file, a named owner, and whatever history your MES or paper travelers can produce; imperfect and written down beats perfect and remembered. Without it, six months from now every conversation about whether the system helped becomes a memory contest, and the person with the most confidence wins.

Run it read-only for a full production cycle. The length is what people cheat on. A full cycle means the ugly shifts are inside it: the changeover, the lot that came in off, the week you ran short-handed, any seasonal behavior your process has. A model that looks excellent across four clean days has told you nothing, because four clean days is not what your plant is. The model makes its call, the existing control makes the real one, and you compare. Log every disagreement with enough context to review later — the disagreements are the dataset. A pilot that is allowed to change the process is not a pilot. It is an uncontrolled deployment with a friendly name.

Write the acceptance number down before you look. This sounds like paperwork and it is the whole thing. Scrap under this. First-pass yield over that. No more than this many interventions per shift. Put it in a file with a date and a name on it before the run starts. When the result lands a little short — and it will — the argument is about whether to change the system, not about whether the bar was ever really the bar. Set the threshold after seeing the output and you didn't run a test; you ran a demo and then wrote a justification for it.

The boundary follows from the rig

Read-only until it earns write access. Anything hard to reverse in physical space stays human-gated until there is a specific reason to move it, and “the quarter is ending” is not a reason. A bad recommendation costs an argument; a bad actuation costs a tool, a batch, or somebody's hand. Not the same risk; not the same permission model.

Autonomy also expands one workflow at a time, not one system at a time. A model that earned write access on setpoint recommendations for one line and one product family has earned nothing anywhere else — every workflow has its own failure mode, and running is the only way you find them. Trust accumulates slowly and it is specific. A plant-wide autonomy tier on a vendor roadmap is a number somebody made up.

Count the net honestly

Net gain gets counted after scrap, downtime, rework, and the hours somebody spends untangling a decision the system got wrong. That last bucket is real labor and it rarely appears in a proposal. If a system saves four hours of scheduling a week and creates three hours of reconciliation, you bought an hour and something new to maintain.

So, the test I would apply before signing anything: if building the measurement is too expensive to justify, the deployment is too expensive to run — you just haven't put a number on the second half yet. Instrumenting the process, pulling the baseline, and running a shadow cycle is a cost you can estimate this week. If that estimate makes the project look bad, the project was already bad and the measurement told you early. That is the cheapest bad news you will ever buy.

What I don't know

I don't know where the ceiling is on closed-loop physical control. My guess is that it is higher than the cautious people say and further out than the demo reels suggest, and I would like NIST or an industry consortium to make this measurable so plants stop improvising a validation method one at a time in the dark.

But it doesn't exist yet, and you are being asked to decide now. Until somebody builds the evaluation, you are it.

Full episode transcript

Nobody is coming to benchmark your production line for you. There is no outside body sitting between a model and your press, your oven, your coating line. There is no independent evaluation a vendor has to pass before they hand you something that moves a physical asset. That job is unclaimed. And if you don't do it, it doesn't get done — the system just goes in, and everybody finds out together. So here's my position: for your physical process, you are the evaluation body, whether you act like one or not. I want to be careful here, because this isn't a complaint that nobody is doing serious evaluation work. They are. On July 27th, NIST announced its Artificial Intelligence Technology Evaluation program — a sequestered testbed with blind tasks, and the first domains are quantum science, genomics, and public safety imaging. That is real methodological rigor and I'm glad it exists. Now look at that list and notice what isn't on it. A stamping press. A curing oven. A coating line running a material that behaves differently in August than it does in February. And here's what convinced me this is structural and not a scheduling problem. NIST ran an AI for Manufacturing workshop in May, and among the topics it listed as areas of interest were the development and validation of AI-driven operations, maintenance, or control strategies, and the measurement and evaluation of the impact of AI on physical processes and assets. Read that as a practitioner. Those are open questions. Not settled methods, not a standard you can hand your quality manager — open measurement-science questions, from the people whose entire job is measurement science. None of that is a knock on NIST. Genomics and public safety imaging are enormous, and you have to start somewhere. What it is, is a map of who's responsible. There's a general-purpose evaluation — can this model do this class of task at all, scored against a held-out set somebody else built. And there's your process acceptance test — will this specific system, on my line, with my material, on my worst shift, make parts I can ship. Two different objects. Only one of them decides whether the line runs, and it isn't the one with the leaderboard. Which brings me to the word I'd like people to stop accepting at face value. Validated. When somebody says a system is validated, that sentence is missing its second half. Validated against what. Their cell. Their material. Their assumed variance, which is usually a distribution somebody typed into a config file. Nobody outside your building has your changeovers. Nobody has your second shift in July. Nobody has the stretch where a supplier quietly changed a coating and your scrap moved before anyone noticed. So make them finish the sentence, out loud, in the room. So build the rig yourself. And I want to be unglamorous about it, because the reason people skip this isn't that it's hard — it's that it's boring and it delays the interesting part. Three pieces. Freeze a baseline of how the process performs today. Run the model against a simulation, or read-only in shadow alongside the existing control, for a full production cycle. Write down the acceptance number before you look at the results. That's the whole rig. It is not sophisticated and it does not need to be. Freeze the baseline first, and freeze it before the vendor is in the building. Cycle time, scrap, rework, unplanned downtime, changeover time — and the human hours nobody logs, the ones somebody spends babysitting the process when it drifts. That last one always gets left out, and it's usually where the real gain is, or isn't. If it isn't written down and dated, then six months from now every conversation about whether the system helped becomes a memory contest, and the person with the most confidence wins that. Then the run, and the length of it is the part everybody cheats on. A full production cycle means the ugly shifts are inside it. The changeover. The lot that came in a little off. The week you ran short-handed. A model that looks great across four clean days has told you nothing, because four clean days is not what your plant is. And it runs read-only next to the existing control the whole time — it makes its call, the control system makes the real one, and you compare. A pilot that's allowed to change the process isn't a pilot. It's an uncontrolled deployment with a friendly name. And write the acceptance number down before you look. This sounds like paperwork and it is actually the entire thing. If you decide what good looks like after you've seen the output, you didn't run a test — you ran a demo and then wrote a justification for it. Pick the number in advance. Scrap under this. First-pass yield over that. No more than this many interventions a shift. Put it in a file with a date and a name on it. Then when the result lands a little short, and it will, the argument is about the system, not about whether the bar was ever really the bar. The approval boundary falls out of the rig, and it does not come from a vendor's autonomy roadmap. Read-only until it earns write access. Anything hard to reverse in physical space stays human-gated until you have a specific reason to move it, and the reason is never that the quarter is ending. A bad recommendation costs you an argument. A bad actuation costs you a tool, a batch, or somebody's hand. Those are not the same risk and they should not share a permission model. And autonomy expands one workflow at a time, not one system at a time. A model that earned write access on setpoint recommendations for one line, one product family, has not earned anything anywhere else. That isn't caution for its own sake. Every workflow has its own failure mode, and the only way you find them is by running. Trust here accumulates slowly and it is specific, and anybody selling you a plant-wide autonomy tier is selling you a number they made up. Now the cost line, plainly, because this is where these projects quietly go negative. Net gain gets counted after scrap, after downtime, after rework, and after the hours somebody spends untangling a decision the system got wrong. That last bucket is real labor and it almost never appears in the business case. If a system saves four hours of scheduling a week and creates three hours of somebody reconciling what it did, you bought an hour, and you also bought something new to maintain. Here's the test I'd apply before signing anything. If building the measurement is too expensive to justify, then the deployment is too expensive to run — you just haven't put a number on the second half yet. That's not a rhetorical trick. Instrumenting the process, pulling the baseline, running a shadow cycle: that is a real cost you can estimate this week. If that estimate makes the project look bad, the project was already bad, and the measurement told you early. That is the cheapest bad news you will ever buy. I'll close honest, because I don't want to pretend I know more than I do. I don't know where the ceiling is on closed-loop physical control. My guess is that it's higher than the cautious people say and further out than the demo reels suggest, and I'd genuinely like NIST or a consortium or somebody to make this measurable, so plants stop improvising a validation method one at a time in the dark. That would be good for everyone, vendors included. But it doesn't exist yet, and you're being asked to decide now. So until somebody builds the evaluation, you're it. Don't let a demo on somebody else's floor be the last thing standing between a model and your production line.

Sources

  • https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite
  • https://www.nist.gov/news-events/events/2026/05/artificial-intelligence-ai-manufacturing-workshop