The embodiment gap.
Data collected on one robot does not transfer cleanly to another. The way a capture programme handles that fact decides how much of its data survives a change of hardware.

Physical AI is embodied by definition, and embodiments are not standardised. Robots differ in their degrees of freedom, in their sensor configurations, in their control interfaces, and in the basic geometry of how they reach, grip, and move. A motor trace recorded on one arm is not a motor trace another arm can follow. Data that is precise and valuable for one machine can be partly or wholly inapplicable to the next, even when the task is identical.
This is the quiet limit beneath the industrialisation of data collection. The move toward large, centralised capture, standardised pipelines, replicated environments, is real progress, but it does not dissolve the embodiment problem. A factory that produces a great deal of data on one configuration produces a great deal of data bound to that configuration. The harder question sits underneath the volume: is a programme building intelligence that generalises across bodies, or simply scaling data tied to a single one.
The distinction matters because it determines what a dataset is worth over time. Data captured without regard to embodiment ages the moment the hardware changes, and hardware changes constantly across a field still converging on its forms. The most reusable data is collected and structured with the embodiment question in view from the start: human, first-person data that captures intent and interaction at a level above any specific machine, and robot-native data recorded with the configuration documented precisely enough that its applicability to another is a known quantity rather than a hopeful assumption.
There is a layer of the pyramid that travels better than the rest. First-person human data captures the structure of a task — the sequence, the contact, the relationship between hand and object — in a form that translates across machines far more naturally than one robot's motor commands translate to another's. Consider the same pick task performed by an anthropomorphic humanoid hand, a parallel-jaw industrial gripper, and a suction gripper: the high-level task semantics may transfer, but the low-level motor actions, grasp geometry, force profiles and trajectories may not. It is not a complete substitute for embodiment-specific data, but it is the part of the dataset least exposed to a change of hardware, which is exactly why its share of a training mix is a strategic choice rather than an incidental one.
This shapes how a programme is built with a lab. Capture is structured around the embodiment a model actually runs on, with robot-native collection matched to that configuration and documented so its reach is explicit, while the broad human layer is weighted to carry the transferable understanding that survives a hardware change. The result is a dataset whose value is understood in advance: what is specific to this machine, what generalises beyond it, and how much of the programme's investment is protected when the embodiment evolves.
The embodiment gap is not a problem that disappears with scale, and treating data as if one robot's experience is every robot's is how a programme discovers, late, that much of what it collected does not transfer. Building with the gap in view, capturing what generalises deliberately and documenting what does not, is how a dataset stays an asset rather than becoming a record of a machine that no longer exists.
— Sources & technical references
— Further reading