Physical AI Reality Check

The three numbers humanoid demos never give you

Industrial buyers decide on cycle time, availability, and MTBF. Humanoid demos publish none of them — and the semiconductor industry has had a standard for two of the three since 1986.

A humanoid robot folds a shirt. It sorts a bin of unfamiliar parts. It walks across gravel without falling. The clip is ninety seconds long, beautifully shot, and it tells you almost nothing about whether the machine can do a job.

That isn’t a criticism of the engineering — those demos represent hard, real work. It’s a criticism of what we’re invited to conclude from them, because the people who actually buy automation decide on three numbers, and demos publish none of them.

Cycle time, measured ready-state to ready-state — the approach, the grasp, the transfer, the release, and the return, not the one segment that got filmed.

Availability: of the hours you wanted the machine working, what fraction it could actually work. A machine waiting for a part or an operator is not producing.

Mean time between failures — plus the half everyone forgets, how long it takes to get running again.

An industry already standardised this

Chip fabs buy capital equipment on long contracts, and they needed to stop arguing with vendors about what “reliable” meant. SEMI E10 — the specification for equipment reliability, availability and maintainability — was “originally published in 1986” and is “one of the most widely used SEMI standards.” Its companion, SEMI E79, covers productivity: overall equipment efficiency and throughput. E10 feeds it, “providing critical equipment time-in-state information used in equipment productivity (OEE) metrics.”

E10’s central move is simple: it establishes six basic equipment states, and every hour of a machine’s life lands in exactly one. Those states are non-scheduled time, unscheduled downtime, scheduled downtime, engineering, standby, and productive.

Notice what that forces open. Standby — the machine is fine, but there’s nothing to run or nobody to set it up. Engineering time — running, but on qualification rather than product. Neither is a failure. Both are hours you paid for and got nothing from.

Non-scheduledUnscheduleddowntimeScheduleddowntimeEngineeringStandbyProductiveWhat the demo showsWhat it leaves out
The six equipment states SEMI E10 defines. Every hour a machine exists lands in exactly one of them. Blocks are illustrative, not to scale — the standard defines the categories, not their proportions.

A demo video, mapped onto E10, is a recording of productive time with the other five states edited out. That’s not dishonest. It just isn’t what a buyer needs — and SEMI is explicit that these metrics exist to set “equipment performance requirements during purchase and service negotiations.”

Why high MTBF numbers mislead

Industrial robot makers publish extraordinary reliability figures: MTBF claims of 40,000, 60,000, 80,000, even 100,000 hours. At face value, 100,000 hours is more than eleven years of continuous operation before a failure.

That number describes the arm, not the cell it sits in — and the cell is what stops.

The resolution is that the robot usually isn’t what fails. Citing the International Journal of Performability Engineering and a survey of 400 factory owners, one automation vendor reports that 80% of failures are not related to the robot — the same finding reported in trade coverage. Both trace to that one survey, so count it as a single data point — and one drawn from conventional industrial robot cells, not humanoids. What it describes is where the failures sit: grippers, fixtures, conveyors, sensors, part presentation, the software gluing it together, and the people feeding it.

Carrying that split over to humanoids is my inference, not the survey’s finding. It rests on a claim you can check independently: the parts that fail are the ones a humanoid still needs.

The number that matters: a 100,000-hour MTBF describes the arm. In that survey, 80% of failures had nothing to do with the robot.

The spec sheet covers the part that rarely breaks. Everything that does — grippers, fixtures, sensors, part presentation, software, people — never appears on it.

The humanoid version is worse, not better

The pitch for general-purpose humanoids is that they delete integration cost. No bespoke cell, no fixtures, no feeders — buy a machine shaped like a person, put it where a person stood.

That aims at the right target. Four failures in five happen somewhere other than the robot, and most of that is the surface a humanoid claims to erase. The question is whether a humanoid removes that surface or merely moves it — because a general-purpose machine relocates the same complexity into perception and control, where failures are harder to predict and much harder to bound. My read: a fixture either holds the part or it doesn’t, and you find out immediately, whereas a vision-and-policy stack that works 95% of the time fails depending on lighting, wear, and whatever the last shift left on the table. Those are different problems to bound, and only one of them has a spec.

And this generation isn’t running unattended. A teleoperation provider reports that models scoring 95% on a benchmark land closer to 60–80% on real tasks. Treat that carefully: the company sells the software and operator networks that fill exactly that gap, and it cites no dataset. But the company exists at all because someone is buying managed operator time, and that is itself a fact about how these fleets run today: someone is watching, and sometimes taking over.

In E10 terms, every intervention is time the machine wasn’t productive. And the labour it consumes is precisely the labour it was bought to replace.

What this means for the money

Whatever payback model a buyer uses, it runs on throughput per dollar — which is why these metrics show up in procurement at all.

Throughput is not the cycle time in the video. Units per hour is availability divided by cycle time, and the division is what does the damage. A 30-second cycle at 50% availability delivers one unit every 60 seconds of wall-clock time. A slower machine — 40-second cycle — at 90% delivers one every 44 seconds. The slower machine wins by a third, and no demo would show you why.

Then there’s the operator ratio. One supervisor to ten machines is a labour-arbitrage business. Closer to one-to-three and you have a business that relocates labour rather than removing it — still potentially valuable, but a different margin structure and a different addressable market. Which one a given company turns out to be is the whole question, and it isn’t extractable from a video.

The checklist

  1. Cycle time, ready-state to ready-state, on a defined task.
  2. Availability over a stated window — ideally the six E10 states, at minimum productive time versus everything else.
  3. MTBF of the deployed system, not the arm, plus mean time to repair.
  4. And one this generation specifically owes you: interventions per hour, and the operator-to-robot ratio the deployment actually runs at.

None of these are unreasonable. Two have had formal definitions since 1986, and a company running a real multi-month pilot already has the numbers — because its customer demanded them before signing.

Where I could be wrong

If the integration surface genuinely collapses — if a machine you retrain by demonstration makes fixtures and feeders obsolete rather than merely different — the argument weakens considerably. That’s the actual bet, and it isn’t absurd. It just hasn’t been shown at production duty cycles.

And if the first wave targets work nobody is currently doing, availability stops being the binding constraint. A robot at 50% isn’t competing with a human at 95%; it’s competing with nothing, and the arithmetic changes entirely.

Both are testable. Neither is testable from a demo reel. The next time one goes viral, the question isn’t whether the robot is impressive — it’s which of the six states you’re being shown, and which five you aren’t.