PHYSICORE
Perspective · 2026

The data nobody is collecting yet.

Pre-training is the land grab everyone can see. The harder, stickier value sits in the data that matches deployment, and almost no one is built to collect it.

A robot working in a real deployment environment that doubles as a data-capture setting.

The race in physical AI data is visible and crowded: vast volumes of broad, first-person data to pre-train foundation models, gathered by anyone with cameras and a workforce. It is the obvious move, and it is already oversupplied. The less obvious work, the work that compounds, sits on the other side of the model. It is the data that matches the exact conditions a system will be deployed into, and the data a system generates once it is working, and it is barely being collected at all.

The reason is that this data is harder in every dimension that matters. Pre-training data rewards breadth and tolerates imperfection. Post-training and deployment data demand specificity: the precise environment a robot will work in, the exact tasks it must perform there, the teleoperation setups matched to that context, and the failures that only appear once a system meets the real conditions of its job. It cannot be gathered speculatively in bulk. It has to be collected deliberately, against a particular deployment, which is why it remains comparatively underdeveloped relative to broad pre-training collection.

It is also where the durable value lives. Pre-training data is a one-time sale into a crowded market; the relationship can end when the training run does. Deployment-matched data is continuous, because a system in production keeps meeting new conditions, and the data that addresses them keeps being needed. This is the difference between supplying a model once and becoming part of how a model improves over its life. The first is a transaction. The second is infrastructure.

The robotics flywheel everyone anticipates depends on exactly this layer, and it does not yet turn on its own. The promise is that once robots are deployed at scale, they generate the data that trains their successors, and the cost of data falls as the value the robots create subsidises it. But that flywheel requires deployed robots at scale to exist first, and they do not yet. Until they do, the deployment-shaped data has to come from somewhere, captured deliberately by people who can operate in the real environments a system is bound for. The flywheel has to be primed before it can spin.

This is the layer Physicore is built to own alongside a lab, not just the pre-training base beneath it. The loop is explicit: deployment failure surfaces the gap; diagnosis identifies the missing data; targeted capture creates it; retraining absorbs it; re-evaluation proves whether performance held. A programme extends past the initial training set into capture matched to the specific deployment: the environments, the tasks, the teleoperation contexts, and the failure conditions a model meets when it leaves the lab, collected and structured so they feed directly into post-training. As the model moves toward production, the capture moves with it, shaped continuously by where the system is actually weak. The pre-training data wins the engagement. The deployment data is what makes it last.

The data nobody is collecting yet is not a niche above the volume business; it is the part of the value chain that is hardest to reach and hardest to displace. A programme that stops at pre-training has sold a commodity. One built to follow a model into deployment, capturing the specific reality it has to survive, becomes the thing a lab cannot easily replace, which is the only position in this market worth holding.