PHYSICORE
Perspective · 2026

The lab is not the world.

Illustratively, a policy that scores strongly on the bench can fall sharply on a real floor. The model has not changed. The world it was raised on was too small.

A robot arm and operator working in a real, cluttered factory floor.

Move the camera six inches to the left, and a policy that succeeded a hundred times in a row begins to miss. The light falls differently, the background is busier, the object sits at an angle it never practised. Nothing about the model has changed. Its world has, and the model was only ever as good as the world it was shown.

This is the most expensive surprise in physical AI, and it is structural rather than incidental. Research success is measured against a test set drawn from the same conditions as training. Deployment, by definition, presents conditions the model was never trained on. Illustratively, a policy that reaches high performance on the bench can drop sharply in a live building, and the fall is not random. It concentrates exactly where reality diverges from the training distribution: the unfamiliar lighting, the cluttered shelf, the worn or mislabelled part, the person who moves without warning.

The arithmetic is what makes this unforgiving. A system running thousands of cycles a day at high but imperfect accuracy still fails many times daily, and each failure pulls a human back into the loop to clear a jam, recover a dropped object, or restart the line. Production economics do not survive that. Real deployments are built around reliability well past ninety-nine percent, and the last few points are the hardest the field has faced, because they live entirely in the long tail of situations no benchmark contains and no clean dataset describes.

More of the same data does not close the gap. A million further demonstrations from a single environment only teach the model that one environment more completely. What closes the gap is data shaped like the place the system will actually work: the same quality of light, the same density of clutter, the same imperfect objects and human interruptions, and above all the same failure modes, captured before deployment rather than discovered after it. The objective is a model that does not encounter reality for the first time on the day it ships.

This is the point at which capture stops being a volume exercise and becomes a targeting problem, which is precisely where the work with a lab begins. A programme starts from the deployment itself: the exact environments a model is bound for, the embodiment it runs on, and the failures already eroding its reliability. Collection is then directed at those conditions, across real sites and the geographies where they genuinely occur, with the difficult cases captured on purpose rather than avoided. The data returns synchronised and labelled in the form the model trains on, and each round is steered by where the previous round moved the policy's performance and where it did not.

That is the difference between selling hours and closing a deployment gap. The lab is not the world, and no amount of bench performance proves a model will hold once it leaves. Bringing the world to the model before the model meets it, in the precise shape the deployment demands, is how the distance between the benchmark and the floor is actually crossed.