Video

Nobody is coming to benchmark your production line for you

NIST stood up a serious sequestered AI evaluation testbed last month and the first domains are quantum science, genomics, and public safety imaging — your press, your oven, your coating line are not on the list, and NIST's own manufacturing workshop still files validating AI-driven control strategies under open measurement questions. That is not a complaint about NIST; it is a map of who is responsible, and the answer is you. Build the rig yourself: a frozen baseline, a shadow or simulated run across a full production cycle, and an acceptance number written down before you look at the results.

What this video covers

  • A general-purpose benchmark and your process acceptance test are different objects.
  • “Validated” is a sentence missing its second half — validated against whose material?
  • Freeze the baseline before the vendor is in the building, and include the unlogged hours.
  • A pilot allowed to change the process is an uncontrolled deployment with a friendly name.
  • Write the acceptance number down before you look at the output.

Watch on YouTube →

Read the companion article →

Full episode transcript

Nobody is coming to benchmark your production line for you. There is no outside body sitting between a model and your press, your oven, your coating line. There is no independent evaluation a vendor has to pass before they hand you something that moves a physical asset. That job is unclaimed. And if you don't do it, it doesn't get done — the system just goes in, and everybody finds out together. So here's my position: for your physical process, you are the evaluation body, whether you act like one or not. I want to be careful here, because this isn't a complaint that nobody is doing serious evaluation work. They are. On July 27th, NIST announced its Artificial Intelligence Technology Evaluation program — a sequestered testbed with blind tasks, and the first domains are quantum science, genomics, and public safety imaging. That is real methodological rigor and I'm glad it exists. Now look at that list and notice what isn't on it. A stamping press. A curing oven. A coating line running a material that behaves differently in August than it does in February. And here's what convinced me this is structural and not a scheduling problem. NIST ran an AI for Manufacturing workshop in May, and among the topics it listed as areas of interest were the development and validation of AI-driven operations, maintenance, or control strategies, and the measurement and evaluation of the impact of AI on physical processes and assets. Read that as a practitioner. Those are open questions. Not settled methods, not a standard you can hand your quality manager — open measurement-science questions, from the people whose entire job is measurement science. None of that is a knock on NIST. Genomics and public safety imaging are enormous, and you have to start somewhere. What it is, is a map of who's responsible. There's a general-purpose evaluation — can this model do this class of task at all, scored against a held-out set somebody else built. And there's your process acceptance test — will this specific system, on my line, with my material, on my worst shift, make parts I can ship. Two different objects. Only one of them decides whether the line runs, and it isn't the one with the leaderboard. Which brings me to the word I'd like people to stop accepting at face value. Validated. When somebody says a system is validated, that sentence is missing its second half. Validated against what. Their cell. Their material. Their assumed variance, which is usually a distribution somebody typed into a config file. Nobody outside your building has your changeovers. Nobody has your second shift in July. Nobody has the stretch where a supplier quietly changed a coating and your scrap moved before anyone noticed. So make them finish the sentence, out loud, in the room. So build the rig yourself. And I want to be unglamorous about it, because the reason people skip this isn't that it's hard — it's that it's boring and it delays the interesting part. Three pieces. Freeze a baseline of how the process performs today. Run the model against a simulation, or read-only in shadow alongside the existing control, for a full production cycle. Write down the acceptance number before you look at the results. That's the whole rig. It is not sophisticated and it does not need to be. Freeze the baseline first, and freeze it before the vendor is in the building. Cycle time, scrap, rework, unplanned downtime, changeover time — and the human hours nobody logs, the ones somebody spends babysitting the process when it drifts. That last one always gets left out, and it's usually where the real gain is, or isn't. If it isn't written down and dated, then six months from now every conversation about whether the system helped becomes a memory contest, and the person with the most confidence wins that. Then the run, and the length of it is the part everybody cheats on. A full production cycle means the ugly shifts are inside it. The changeover. The lot that came in a little off. The week you ran short-handed. A model that looks great across four clean days has told you nothing, because four clean days is not what your plant is. And it runs read-only next to the existing control the whole time — it makes its call, the control system makes the real one, and you compare. A pilot that's allowed to change the process isn't a pilot. It's an uncontrolled deployment with a friendly name. And write the acceptance number down before you look. This sounds like paperwork and it is actually the entire thing. If you decide what good looks like after you've seen the output, you didn't run a test — you ran a demo and then wrote a justification for it. Pick the number in advance. Scrap under this. First-pass yield over that. No more than this many interventions a shift. Put it in a file with a date and a name on it. Then when the result lands a little short, and it will, the argument is about the system, not about whether the bar was ever really the bar. The approval boundary falls out of the rig, and it does not come from a vendor's autonomy roadmap. Read-only until it earns write access. Anything hard to reverse in physical space stays human-gated until you have a specific reason to move it, and the reason is never that the quarter is ending. A bad recommendation costs you an argument. A bad actuation costs you a tool, a batch, or somebody's hand. Those are not the same risk and they should not share a permission model. And autonomy expands one workflow at a time, not one system at a time. A model that earned write access on setpoint recommendations for one line, one product family, has not earned anything anywhere else. That isn't caution for its own sake. Every workflow has its own failure mode, and the only way you find them is by running. Trust here accumulates slowly and it is specific, and anybody selling you a plant-wide autonomy tier is selling you a number they made up. Now the cost line, plainly, because this is where these projects quietly go negative. Net gain gets counted after scrap, after downtime, after rework, and after the hours somebody spends untangling a decision the system got wrong. That last bucket is real labor and it almost never appears in the business case. If a system saves four hours of scheduling a week and creates three hours of somebody reconciling what it did, you bought an hour, and you also bought something new to maintain. Here's the test I'd apply before signing anything. If building the measurement is too expensive to justify, then the deployment is too expensive to run — you just haven't put a number on the second half yet. That's not a rhetorical trick. Instrumenting the process, pulling the baseline, running a shadow cycle: that is a real cost you can estimate this week. If that estimate makes the project look bad, the project was already bad, and the measurement told you early. That is the cheapest bad news you will ever buy. I'll close honest, because I don't want to pretend I know more than I do. I don't know where the ceiling is on closed-loop physical control. My guess is that it's higher than the cautious people say and further out than the demo reels suggest, and I'd genuinely like NIST or a consortium or somebody to make this measurable, so plants stop improvising a validation method one at a time in the dark. That would be good for everyone, vendors included. But it doesn't exist yet, and you're being asked to decide now. So until somebody builds the evaluation, you're it. Don't let a demo on somebody else's floor be the last thing standing between a model and your production line.