NVIDIA’s video-generation lab has added robot control to its LongLive repository. The new project, Long-WAM, trains and evaluates robot policies that draw on a long visual memory, and ships interfaces for three named machines: the YAM bimanual arm, the Franka arm, and the Unitree G1 humanoid.
The repository changelog dates the Long-WAM release to October 7 and the accompanying arXiv paper to October 8. Until now the NVlabs/LongLive repository was a single video model, and our earlier coverage described its LongLive 2.0 release. It is now an umbrella over four self-contained projects, and the README is built so that a reader can fetch only the one they need.
A world-action model, in the README’s usage, sits between video generation and robot control. It learns from a long run of what a robot has seen and uses that to choose what the robot does next. The Long-WAM directory covers training, evaluation, generating robot video, and deployment. Its listed benchmarks are LIBERO, RoboTwin 2.0, Domino, RoboCasa GR1 and RoboCasa365, and its runtime profiles target an RTX 5090, a DGX Spark and an AGX Thor.
The checkpoint list shows what “long context” means in practice. One RoboCasa GR1 family comes in five variants, from zero seconds of context up to 19.2 seconds, doubling at each step after 2.4. That is a built-in experiment on how much memory a policy needs. A separate YAM checkpoint is described as bimanual pretraining on a dataset the README calls ABC 130K, and a RoboCasa365 model is labelled a generalist kitchen policy.
Why does a video lab end up writing robot policies? Both problems fail the same way. A video model that forgets what it generated thirty seconds ago produces a room whose furniture drifts. A robot that forgets what it did thirty seconds ago picks up the cup it already stacked. The LongLive line of work is about keeping generated video coherent over long sequences, and Long-WAM, which the README says is built on LongLive 2.0, applies that machinery to the harder case where the output moves a physical arm.
The other three projects are older or sideways steps. LongLive 2.0 handles training and serving with NVFP4 (NVIDIA’s 4-bit number format, which shrinks memory use by storing values in fewer bits) and sequence parallelism. LongLive-Plug is a distillation method, meaning it transfers a learned capability from one model to another, here trained once on a base model and reattached to compatible downstream models without retraining. It supports MiniMax-H3 and two Wan video models: the 14-billion-parameter Wan2.1-14B, and Wan2.2-TI2V-5B, the 5-billion-parameter one. LongLive 1.0 is the original interactive video generator.
The repository’s own labels give a rough maturity ladder. Long-WAM, LongLive-Plug and LongLive 2.0 each point to an arXiv preprint, which means the paper has been posted but not peer reviewed. Only LongLive 1.0 carries a conference venue, ICLR 2026, where reviewers vetted it. The newest work is the least checked.
On licensing, the README says the whole repository is released under Apache 2.0, with a copy of the licence and any third-party notices inside each directory. That covers the code; the README also says each project has its own model weights, but does not separately state their terms, so check the checkpoint pages before commercial use.
The README quotes no Long-WAM task success rates in the text we reviewed. Those numbers live in the paper, and they are NVIDIA’s own results on benchmarks the same team chose to report.
For anyone building on a Franka or YAM arm, the practical move is to pull the Long-WAM directory alone and run the released LIBERO or RoboTwin checkpoints against your own task before reading anything into the headline benchmarks.
NVIDIA’s NVlabs LongLive repository README on GitHub (github.com/NVlabs/LongLive), with release dates taken from its changelog.