PHYSICORE
Perspective · 2026

Provenance decides what you can build on.

As physical AI moves into homes, hospitals, and factories, the question stops being how much data a model was trained on. It becomes whether anyone can prove where that data came from.

A multi-camera capture rig in a real clinical environment, showing a recording session with session metadata and chain-of-custody details.

For most of the last decade, training data was treated as something to be acquired by whatever means worked, with provenance an afterthought. That era is closing. As physical AI moves out of research and into homes, hospitals, factories, and public space, the systems trained on this data will operate around people, under regulation, and inside enterprises that will not deploy what they cannot account for. The question shifts from how much data a model has seen to whether its origin can be proven, and the second question is far harder to answer after the fact.

The exposure is structural, not cosmetic. A model trained on data of uncertain origin carries that uncertainty into every deployment built on it. If the rights to use the data were never clear, the risk does not stay with the collector; it travels downstream to the lab that trained on it and the enterprise that deployed it. Footage gathered without a clean chain of consent and rights is not a discounted asset. It is a liability embedded in the model, invisible until the moment it matters, and difficult to remove without retraining.

This is why provenance cannot be retrofitted. The clean record of where data came from, under what consent, with what rights to use and resell, has to be established at the moment of capture. It cannot be reconstructed later for a pile of footage of unknown origin, and a dataset assembled without it cannot be made safe by relabelling. The integrity is either built in from the first frame or it is permanently absent, which means the discipline of provenance is a property of how an organisation collects, not a document it produces afterward.

It is also becoming the constraint that decides which suppliers survive. As labs sign exclusive licences and harden their data relationships, the field is consolidating from many vendors to few, and the ones that endure will be those whose data is not only useful but defensible, owned at the source, cleared for use, traceable end to end. Only data that can prove its origin will clear the bar for systems deployed into regulated reality. The market is moving from how much, to how much you can actually build on.

This is ground Physicore was built to hold, and it is where an evidence-grade pedigree becomes structural rather than incidental. Data is captured with consent, rights, and chain of custody established at the source and carried through every stage, so what reaches a lab is rights-cleared, traceable and auditable, with the provenance intact rather than asserted. Provenance supports enterprise, legal and regulatory due diligence — it does not by itself make a model legally or regulatorily deployable, but it is a precondition for either. A model trained on data whose origin can be answered enters regulated environments without carrying an unanswerable question about where its training data came from. The rights are not a clause added at the end. They are a property of the data itself.

Provenance is not a feature, and treating it as one is how a programme discovers, too late, that the data underneath its best model cannot be defended. In a field about to move into the most sensitive environments there are, the data that can prove what it is becomes the only data worth building on. Everything in this series points to the same conclusion: the volume was never the hard part, and it was never the valuable part. What a model can actually be built on is.