Robotics startup Dyna Robotics has published a technical report on Dyna-2, a model it calls a “world-action model,” trained on upward of a million hours of first-person human video footage. The company reports that prediction accuracy on held-out data improves in a smooth, predictable pattern as the video pile grows, from one thousand hours up through one million. That pattern is a scaling law: the same kind of curve that justified pouring billions into large language models, on the logic that more of the raw material buys a proportional, forecastable return.

A world-action model works differently from the vision-language-action systems most robotics labs favor. A VLA maps a camera image and a language instruction straight to a robot command. Dyna-2 instead learns to predict what a scene will look like a moment later and generates the action alongside that prediction, sharing one network trunk between the two objectives. The video-prediction half of that task needs no robot hardware at all. It trains on ordinary head-mounted footage of people cooking, folding laundry, and assembling parts, footage that already exists in vast supply. That is why video hours, not robot-collected demonstrations, become the unit Dyna Robotics scales: video is cheap and plentiful, while paired robot-action data is not.

The company’s central claim goes further than a clean curve on human footage. It reports that the same checkpoints, trained purely on human video with zero robot data at any stage, were then scored on 39 robot manipulation tasks they had never encountered, and accuracy still climbed in step with the volume of human footage used during pretraining. Dyna Robotics frames this as the first scaling relationship to bridge human and robot embodiment. Carried through to physical post-training on 14 real tasks, the aggregate score rose from 20 percent to 53 percent of each task’s ceiling as pretraining scaled from a thousand to a million hours. One task, opening a lockbox by turning its key, failed at every budget through 100,000 hours, then worked in roughly 9 of every 10 attempts once pretraining hit a million.

That climb is the interesting part, and also the part nobody outside Dyna Robotics has checked. The 39 held-out robot tasks, the 14 post-training benchmarks, and the head-to-head comparison against the company’s prior model, Dyna-1, were all designed, run, and scored by Dyna Robotics or evaluators it selected. The company reports its new model won 65 percent of head-to-head trials against Dyna-1 and, deployed cold at real customer sites neither model had seen, passed 87 percent of the time versus Dyna-1’s 46 percent. Both are comparisons between two models from the same lab, not results checked against an outside benchmark or reproduced by another research group.

A scaling law is a particular kind of claim: that spending more, here on video collection, buys improvement that is predictable and keeps compounding, which is exactly the argument that turns a data-collection strategy into a fundable pitch. Whether it holds up depends on whether watching humans genuinely teaches a model the physics of contact and grasp, or whether Dyna Robotics built an evaluation set that happens to reward the specific gap it filled with more data. Settling that needs an outside lab running the same human-to-robot transfer test on hardware and tasks Dyna Robotics did not choose, ideally with a held-out set another party selected rather than the company itself.

Robotics teams weighing pretraining recipes should treat this as a hypothesis worth testing on their own rigs, not a settled formula to copy. Anyone pricing a data-collection round pitched on this curve should ask what it looks like on robots and tasks the company had no hand in picking.

Dyna Robotics published these findings in a technical report on its own research site, dyna.co/dyna-2, with no release date given.