PHYSICORE

Perspective · 2026

WHAT GOOD ROBOT DEPLOYMENT LOOKS LIKE.

Reliability is one input among several. What to establish about a robot system before committing capital to a rollout.

Physicore ResearchPublished Last reviewed

Key takeaways

  • Reliability is one input among several. A system can hit its success rate and still fail the business case on cycle time, downtime response, changeover or integration.
  • The most expensive error in robot procurement is accepting a success rate without establishing the conditions under which it was measured.
  • Acceptance criteria should specify conditions, not only outcomes. A completion rate with no stated lighting, clutter, traffic or surface conditions is difficult to enforce and easy to dispute in good faith.
  • Failure classification matters more operationally than the headline number. A system that stops safely can be worked around. A system that fails and continues cannot.
  • Most pilots run under conditions no production line sustains: attentive supervision, a motivated team, a single shift, a stable product mix.
  • The party demonstrating the system should not be the party certifying it.
A robot operating in a live industrial environment, where deployment conditions differ from demonstration conditions.

HOW IT GOES WRONG WHILE EVERYONE BEHAVES REASONABLY.

There is a version of robotics procurement that goes wrong while everyone involved behaves reasonably.

A vendor demonstrates a system. It performs the task repeatedly and well. The success rate is high, the reference sites are real, the team is credible. The pilot runs for eight weeks in a defined area and meets its criteria. The business case is approved on the strength of it.

Then the rollout underperforms, and the reasons are rarely the ones anybody was tracking. Cycle time was measured at the robot rather than across the line either side of it. The pilot ran on day shift with the site's best operators nearby. It ran through a stable product mix, and the seasonal packaging change arrived in November. Downtime response looked adequate while the vendor's engineer was on site twice a week, and looked different when the site was on its own at three in the morning.

None of that requires anyone to have misled anyone. It requires only that the pilot measured something narrower than the thing the business was buying.

WHAT IS ACTUALLY BEING BOUGHT

The system performing a task is one component. The rest determines whether the deployment holds.

Throughput, measured across the line rather than at the robot. A system with a strong task success rate that runs slower than the operation it replaces can still fail on economics. What matters is the effect on total throughput, including what the surrounding process absorbs when the robot stalls, retries or hands back an exception.

Exception rate, and what each exception costs. A five per cent failure rate is a different proposition depending on what happens next. If a failed pick is re-attempted automatically in four seconds, it is noise. If it needs a person to walk over, clear a jam and reset, it is a labour line that was never in the business case.

Downtime response. Time to recovery usually decides whether operations tolerate a system. Who diagnoses a fault overnight, whether the site can clear it without the vendor, what the spares lead time is, and whether the support commitment matches the shift pattern the site actually runs.

Integration. The system connects to a warehouse or manufacturing execution system, and the value depends on that connection being reliable and on someone owning it when it is not. Integration effort and its ongoing maintenance are routinely underestimated at the point of purchase.

Facility readiness. Power, network coverage, floor condition and load, guarding or safety sensing, charging or docking positions, and the floor space taken out of productive use. These appear in a business case only if someone puts them there.

Changeover and product mix. Most pilots run against a stable set of items. Production does not. What happens when packaging is redesigned, a seasonal line arrives, a supplier substitutes a carton, or the mix shifts by twenty per cent is worth answering before the commitment rather than after.

The people around it. The operators working alongside the system, the maintenance team expected to support it, and where applicable the works council or union consultation that has to be done properly. Deployments fail on adoption at least as often as on technology.

Site variance. A system proven at one site is evidence about that site. The second has a different layout, different lighting, different local practice and different people. Multi-site rollouts underperform most often because the first site was treated as a template rather than a sample.

Vendor continuity. A multi-year commitment to a company that may be acquired, change direction, or run short of funding. Escrow, data portability and a defined exit position are inexpensive to agree at the outset and difficult to agree later.

Data rights. The deployment generates footage of your operation, your staff and your processes. Who owns it, what it may be used for, whether it may be used to train models sold to others, and what happens to it when the contract ends. Settle this before installation.

And underneath all of it, reliability under your conditions. This is where the industry's evidence is weakest, and it is the part we work on.

WHY THE RELIABILITY EVIDENCE IS USUALLY WEAK

Not through any intent to mislead. Demonstrations are built to be repeatable, and repeatability requires control: fixed lighting, an uncluttered surface, a known object set, a clear workspace.

Those controls remove the conditions that cause deployment failure. The result is a measurement taken at the most favourable point available and presented as a general property of the system.

The gap is visible at the data level. In our audit of one of the most widely used robot manipulation datasets, one run in a hundred presented any real environmental difficulty and ninety-one in a hundred were brightly and evenly lit. Systems learn from footage of tidy rooms, are evaluated under conditions resembling that footage, and are then deployed into places that are neither.

A success rate presented without its conditions is therefore not a property of the system. It is a property of the system under circumstances that were never stated, which means it cannot be compared between vendors, reproduced at your site, or relied on in a contract.

THE FAILURE CLASSIFICATION THAT MATTERS

Most reporting records whether the task succeeded. Operationally, the more useful distinction is what happened when it did not.

OutcomeOperational consequence
CompleteTask achieved within tolerance.
Complete with faultAchieved, but with unintended contact, excessive retries or a timing overrun. Accumulates as wear, damage and cycle-time drift.
PartialAchieved in part. Usually requires human resolution, and is often absent from vendor reporting entirely.
Failed, detected, safe stateThe system knew it had failed and stopped safely. Recoverable, plannable, insurable.
Failed, undetected, continuedThe system did not know it had failed and carried on. This is the outcome that damages product, equipment and confidence.

A headline success rate treats the last two identically. Operations does not.

WHAT GOOD LOOKS LIKE

StageWhat to establish
Before pilotBaseline performance measured in your conditions, not the vendor's. Written acceptance criteria specifying outcomes and the conditions under which they are measured. Facility readiness costed. Integration scope and ownership agreed.
During pilotStructured variation rather than steady-state running: conditions raised deliberately, one at a time, to locate where performance degrades. Exception handling timed and costed. Downtime response tested without the vendor on site.
Failure handlingFailures classified by the table above. Recovery tested by inducing failure deliberately rather than waiting for it.
Before scalingPerformance re-measured under compound conditions, because deployment presents difficulties in combination. Changeover and product-mix variation tested. Second-site variance assumed rather than hoped against.
CommercialData rights, licence position and provenance of the training data settled in contract. Escrow, portability and exit position agreed. Support commitment matched to the shift pattern the site runs.
EvidenceIndependent verification where the commitment is material or where vendors are being compared.

The phased structure now appearing in large humanoid programmes reflects part of this. Publicly announced deployments have included a capability demonstration period followed by a separate period validating stable, continuous operation approaching production scale, before any expansion. That sequencing exists because pilot performance and production performance are different measurements.

THE QUESTION WORTH ASKING

Not “how good is this system”, which invites a number.

Instead: under which of my conditions does it stop working, how does it fail when it does, and what does that failure cost per shift.

The first part requires testing that raises difficulty deliberately until performance degrades, and reports where the boundary sits. The output is not a score but a profile: performance across each condition, at each intensity, with failure modes and recovery behaviour recorded separately.

The second and third parts are yours, and no external party can answer them for you. But they cannot be answered at all without the first.

HOW THE PARTIES FIT TOGETHER

There are three roles in a deployment, with different interests, which is healthy rather than a problem.

The robot maker or lab wants the system to succeed and improve. Their interest is in knowing precisely which conditions break performance, because that determines what to fix and what data to collect.

The deployer wants certainty before committing capital, in terms that survive contact with their own operation.

Independent evaluation serves both. The same testing that tells an operator whether to proceed tells the maker exactly where the system needs work. A condition profile is a procurement document and a development specification at once, which is why it is worth producing properly, and by a party with no stake in the outcome.

Physicore tests robot systems in the environments they are built for, identifies the conditions under which performance breaks down, and where required builds the data that closes the gap. We sell no robots and build no competing model.

The specification behind the work is published in the Physicore Standard v0.1. The evidence for why it is needed is in what the data actually contains, and the argument for testing outside the laboratory is in a bench is not a building.

FAQ

Is a vendor demonstration not sufficient evidence?
It establishes capability. It does not establish reliability under your conditions, because demonstrations are controlled by design, and control removes the conditions that cause deployment failure.
What should acceptance criteria contain?
The outcome required, the conditions under which it is measured, the exception rate and its handling cost, and the classification of failures as detected or undetected.
Why does the detected and undetected distinction matter?
A system that detects a failure and reaches a safe state can be planned around and insured. A system that fails and continues damages product, equipment and confidence. A headline success rate treats both identically.
Our pilot succeeded. Why would the rollout differ?
Pilots typically run with attentive supervision, a motivated team, one shift pattern and a stable product mix. Production sustains none of those. Test changeover, shift variation and second-site differences before scaling.
When is independent evaluation worth it?
Where the commitment is material, where vendors are being compared, or where the deployment environment differs substantially from the conditions in which the system was demonstrated.
Does it slow the rollout?
Structured variation locates the performance boundary within a defined period. Discovering the same boundary through production experience takes longer and costs more.

This article describes general evaluation and deployment practice and does not constitute procurement, legal, safety or engineering advice. Requirements vary by jurisdiction, sector and application.