Home Top Stories How Operational Insights Are Advancing With Physical AI
Top Stories

How Operational Insights Are Advancing With Physical AI

Share
How Operational Insights Are Advancing With Physical AI
Share

This article was co-written with Gilad Langer, Industry Practice at Tulip.

In 1958, commercial aviation started putting flight data recorders on airplanes. Every accident since has been reconstructed second by second, and an entire safety discipline got built on top of that record. Surgery has been getting operating-room black boxes. Every professional sport now has frame-accurate replay, and most of them have quietly rewritten their own rules based on what the replay turned out to show.

Manufacturing, the sector that actually makes the physical world, still investigates a defect by asking three people what they remember about Tuesday.

That’s our state of the art.

We don’t point that out to be cheeky. We point it out because this industry is about to spend an enormous amount of money teaching AI to help run its operations, and remarkably few people are asking what it’s going to learn from.

The binding constraint on industrial AI isn’t model quality. It’s that the models are being asked to learn from a record that was never a record of reality in the first place. And we don’t think the systems that produce that record, the ones this industry has spent thirty years installing, survive the fix.

Despite being an engineer, I’ve come to believe that most of the excitement about flashy industrial technology is a distraction from the boring thing that actually matters. This is a piece about the boring thing.

Systems Of Assertion

Walk into most plants and you’ll find a stack: an MES, a QMS, a LIMS, a warehouse system, a historian, and some number of homegrown applications holding the seams together. The market treats these as different products, in different categories, from different vendors. Architecturally they’re one system at different scopes, and they rest on the same assumption.

Every one of them records what somebody said happened. An operator confirms a step. A technician enters a result. A supervisor signs off on a deviation. The schema decides in advance what is worth knowing, and then a human being types reality into it.

These are systems of assertion. They’re genuinely good at what. A quality check failed at 10:02. A line went down for eleven minutes. A batch deviated from its recipe. Not one of them can tell you how. In fairness, none of them was ever designed to.

There are two observations about operations that we think are foundational and that this industry keeps designing around instead of designing for:

Every operation is different, and none of them runs the way it’s documented. The process as designed and the process as executed diverge continuously. That divergence is where all the value and all the risk live. It’s also precisely what a predefined schema can’t capture, because you have to know what to ask before you can record the answer.

Human judgment is the most consequential variable on the floor, and the least instrumented one. Manufacturing has telemetry on every motor in the building and almost none on the decisions that actually determine whether the shift goes well.

We spent thirty years instrumenting the machines and almost none instrumenting the work.

What A System Of Observation Changes

If the operational record can be observed rather than asserted, if cameras, sensors, and models that genuinely understand what they’re looking at can produce a continuous, searchable, time-aligned account of what happened, then the schema stops being the thing that defines what’s knowable.

The schema is the product. It’s what you buy. It’s what you configure for eighteen months. It’s what you validate, and what you can’t easily change afterward. A system whose central value proposition is a well-governed place to type things is in an awkward position the moment typing stops being how the record gets made.

What survives is governance, traceability, approval, genealogy, and a defensible audit trail. Nobody serious is arguing those go away. If anything, they get harder and more important, because you’re suddenly governing far more data than before.

What doesn’t survive is the assumption that all of that requires a monolithic, schema-first system of record that people feed by hand.

We’re not predicting that a list of vendors disappears. We’re predicting something more specific and more uncomfortable: that the architecture becomes a liability faster than the installed base turns over. The gap between those two curves is where the next decade of this industry gets decided.

It is also where speed of change stops being an abstraction. A frontline team should be able to see a problem and have a working change running in production in hours to days. Traditional schema-first architectures mean slow changes due to complexity and rigid requirements like change control, revalidation, regression testing, and vendor involvement. Rigidity is a risk.

“The King Is Dead, Long Live The King!”

That’s not our headline. It’s Tom Comstock’s, at LNS Research, from a piece he published in August.

Comstock has been in the MES business for more than forty years, and he reported something that cuts directly against everything above. Every discrete MES vendor LNS surveyed is growing. Roughly 30%-plus compound annual growth over three years. Adoption is increasing. On-premise deployment still dominates, and some cloud-only entrants are now shipping on-prem versions. The product roadmaps are aggressive and well funded.

He’s right about the data, and we’d rather engage it than pretend we didn’t read it.

But growth and fitness aren’t the same measurement. Kodak’s best year for film was 1999. Nokia shipped more phones in 2007 than in any year before it. Revenue inside a large installed base is a lagging indicator of what you already sold, not a leading indicator of whether the architecture still fits the problem.

Two of Comstock’s other findings make the case better than we can make it ourselves.

The first: he notes that some vendors are “trying to make MES composable” by disaggregating products they already had. I’ve written before that a solution isn’t composable if the vendor needs to do the composing. Breaking a monolith into modules that only the vendor can reassemble isn’t composability. It’s a longer statement of work.

The second finding is the one that should give the whole category pause. Comstock’s ninth observation is that MES companies “do not see the need for or demand for” connected frontline worker applications. Not that they tried and struggled, that they don’t see the demand. A category that doesn’t believe the person doing the work is the point is not going to be the category that learns from them.

None of which means any of this happens quickly. These systems run production. You can’t experiment on them. They’re entangled with validated states, regulatory commitments, and twenty years of accumulated customization that nobody fully understands anymore. Brownfield reality is unforgiving. That’s exactly why this is a decade-long shift and not a quarterly one.

The Data Nobody Can Simulate

Gartner’s 2026 assessment of physical AI calls it transformational, and then says in the same breath that its hype and risk currently outweigh its business outcomes, and that CIOs should plan carefully to avoid costly false starts. Both halves of that are true. Anyone selling you only the first half is selling you something.

The physical AI field knows perfectly well that data is its bottleneck. Jensen Huang said it plainly at CES this year: real-world data collection is slow, costly, and never enough. The response has been an enormous investment in synthetic data that simulates the world, then generates the training set, and engineer around the collection problem.

While synthetic data works well for teaching a robot to move, it cannot capture the reality of your actual operations. Simulation can only generate what you’ve already modeled. The thing you actually need to learn is the deviation, the workaround, the improvisation, the fixture somebody modified in 2019 for reasons now lost to the company. Deviations are by definition not in your model. That’s what makes them deviations.

Real operators doing actual work remain the one dataset that cannot be replicated through simulation, specifically, all the ways it departs from the plan. Which means the scarcest input to physical AI in manufacturing isn’t compute, and isn’t model architecture. It’s a real record of what actually happens on your floor, and almost nobody has one.

There’s a version of this idea that Lean people will recognize immediately. Kaizen has always been reinforcement learning from human feedback: observe, hypothesize, change, measure. The loop is a century old. What’s new is that the observation can finally be captured systematically, instead of with a stopwatch, a clipboard, and one industrial engineer’s afternoon.

This reframes the workforce conversation, too. In this telling, the frontline worker isn’t the thing being automated away. The frontline worker is the signal, the source of the only training data that matters. Deloitte and The Manufacturing Institute put the need at 3.8 million US manufacturing jobs through 2033, with as many as 1.9 million potentially going unfilled. Capturing what experienced people know while they’re still in the building stops being a technology question somewhere around there.

Agency Requires Context, Not Only Autonomy

As agents start proliferating in the industrial operations space it is useful to think about three tiers:

The first is reflex. Some things have to happen in under a hundred milliseconds: a vehicle gets too close to a person, a warning sounds. That should never route through an agent, and building it that way is a mistake dressed up as innovation. Determinism wins.

The second is deterministic alerting on well-defined events. Unglamorous, necessary, and probably necessary for far longer than anyone wants to admit.

The third is real agency: something that holds current state, reasons over it, and acts through tools. That’s the interesting tier, and it’s where the over-promising is worst.

True agency is gated by context, not by autonomy. Give an agent a verifiable, real-time picture of what’s actually happening and it becomes useful. However if you only give it a database extract you risk building a liability with a confident tone. The record comes first and the agents come second, which is the reverse of the order most of this market is currently selling.

The Bind In Regulated Manufacturing

The sharpest version of the whole problem shows up where the stakes are highest.

When a batch deviates in a regulated environment, the manufacturer is required to determine root cause. And FDA, EMA, and WHO have all made clear (correctly, in our view) that “human error” is not an acceptable root cause. Investigators get cited when their investigations stop there.

So consider the position this leaves a quality organization in. You’re required to explain what happened. You’re forbidden from blaming the operator. And you’ve been equipped with a record that captures only what somebody asserted after the fact, in fields somebody else chose years ago.

You’ve been handed an obligation and denied the instrument.

Now imagine the record were produced by the work rather than typed alongside it. The batch record, the device history record, the operational history: all of it a byproduct of execution rather than a separate documentation effort layered on top of the actual job. Not less evidence, actually more of it, and better, at a fraction of the effort, because nobody had to stop working in order to write it down.

If that works, it’s the largest prize in this entire category. But let’s be precise about the tense: if. Nothing about this is validated today, and I’d be suspicious of anyone telling you otherwise.

Drawing The Line

GDPR pushes toward anonymizing people in operational data. GxP data integrity demands attribution, think the A in ALCOA. Those pull in opposite directions, and in a regulated environment the regulator wins. Which means any credible design has to be built to survive attribution, not to avoid it. Nobody has solved that.

In addition, the EU AI Act has been unambiguous since February 2025: inferring a worker’s emotional state from biometric data in the workplace is prohibited outright, with narrow medical and safety exceptions, and penalties reaching seven percent of global turnover. Read the Act’s own reasoning: part of the justification is that the technology doesn’t reliably work. That’s a red line we cannot cross.

Observing how work gets done and inferring how a worker feels are not the same thing. But that distinction has to be designed in, not asserted in a press release. In Germany, a works council has co-determination rights over the introduction of technical monitoring. That isn’t an obstacle to route around but rather a design constraint, and a reasonable one.

Which brings us to the principle we’d put above all of this. You cannot only take data out of an operation. You have to put value back in, or the people being observed will correctly conclude that they’re being surveilled, and they’ll be right.

But This Is Just The Beginning…

The realistic status is that this class of technology is early in its development. The temporal reasoning isn’t yet good enough to be trusted on fine-grained manual work in most settings. Consent design that a works council would sign off on doesn’t exist as a standard. Retention and deletion policy is being invented case by case, badly.

You’ll also notice there are no ROI numbers in this article. Honestly because they don’t exist yet, and anybody publishing them today is modeling, not measuring. Apply that test to every claim you hear about this category over the next twelve months, including the ones you hear from me.

None of this gets manufacturing all the way from where it is to grounded physical AI. It isn’t the complete recipe. But we think it’s the missing ingredient, and we think the industry has spent the last few years arguing about the wrong things.

Aviation didn’t get safer because pilots got better. It got safer because we started recording what actually happened, and then built a discipline around learning from it. The recorder came first. The safety culture came second.

So, do you know how your best shift differs from your worst one, or only that it does? When your next deviation happens, will you have evidence, or will you have recollections? And if you’re building an AI strategy on top of systems that can only tell you what somebody typed, what exactly is it learning?

These are the questions a few hundred of us will be arguing about in Somerville in October. On this one, I’d rather be argued with than agreed with.

Source link

Share

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *