Jan Bosch, a professor who advises companies building physical systems, published an essay on his personal site on 7 September 2026 arguing that the industry’s comforting story about embodied AI is wrong. The pitch has been that vision and language are solved problems, so wiring a capable model to an arm and a camera is mostly an integration exercise. Bosch says that framing skips over the part that actually determines whether a robot works.
His case rests on a distinction between where robotics has already succeeded and where it keeps stalling. Fixed-task automation, a robot welding the same joint on the same fixture thousands of times a day, already works at industrial scale. Bosch cites the International Federation of Robotics count of 542,000 industrial robots installed in 2024 and an operational stock of 4.66 million machines worldwide. Those systems hit their reliability targets by refusing to generalize: everything outside a narrow, engineered envelope is designed away rather than learned.
General-purpose manipulation is the opposite story, and Bosch points to a concrete number for why. A language model trains on text that already existed for other purposes and simply had to be gathered. No equivalent corpus exists for the case of a robot lifting an object it has not seen, oriented in a way it has not met. Each example of that has to be physically generated, usually by a human teleoperating a real machine in real time, which means the training set grows only as fast as someone is willing to pay for robot hours. That is the load-bearing constraint in his argument: manipulation data is expensive because it cannot be scraped, only manufactured one motion at a time.
To support it he points to a survey he describes as covering 1,228 vision-language-action papers, spanning February 2023 through June 2026. On the Libero benchmark, he reports, success rates near 95 percent fall below 30 percent once the scene is perturbed even modestly, and on a separate benchmark a model’s score dropped from over 90 percent to exactly zero. He is characterizing someone else’s benchmark aggregation rather than a measurement of his own, and the framing he draws from it, that this looks like memorized trajectories breaking the moment reality drifts outside them, is his interpretation of those figures.
AI Insiders reported on 8 September 2026 on a frontier model given direct control of robot arms: it placed a block in a bowl 19 times out of 20 but seated a puzzle piece correctly only twice in 20. That gap between a well-covered motion and a lightly-covered one is the same failure mode Bosch is describing at benchmark scale, just visible in a single afternoon of testing rather than across 1,228 papers.
From there Bosch draws his advice for founders: stop promising a general-purpose robot and instead pick one task, one environment, one gripper, and collect enough real trajectories to cover that narrow slice completely. He frames the seed-stage question as how many hours of interaction are needed before failure rates become acceptable, and who is paying for them, rather than how capable the underlying model is. The largest unused asset, on his reading, belongs to the incumbents. Motion data pours off those 4.66 million installed machines daily, and hardly any of it is being recorded in a shape that could train anything.
What Bosch does not spell out is why his own advice is hard to act on. A narrow, instrumented deployment is not really a software company, it is a services business with a hardware bill attached: robots, sensors, integration labor and an operator to run it all before a single dollar of recurring revenue shows up. That is close to the exact shape of business that venture capital has spent the past decade steering away from, favoring software margins over equipment-heavy operations. The mismatch shows in the figures he cites. Robotics drew $40.7 billion of venture money during 2025. Against that, BMW’s Figure 02 deployment recorded roughly 1,250 hours of operation across eleven months, which works out at under four hours in a working day.
Bosch’s underlying point is that physical systems tolerate failure differently than text does. A wrong chatbot answer costs a reader a few seconds. Put a humanoid next to a human being and the reliability bar he cites moves above 99.9 percent rather than 95, a jump he treats as a change of engineering discipline rather than a matter of degree.
For operators evaluating a robotics vendor, Bosch’s framing suggests asking for perturbation results rather than a demo reel, since a demo only shows performance inside the training distribution. Any team unable to show what happens when lighting or object position shifts has not measured the number that matters.
Jan Bosch, writing on his personal site on 7 September 2026.