Runway released GWM Worlds 2, a world model that generates interactive video environments in real time, complete with synchronized audio and no fixed session length. A user types actions such as “swing the spear” or “flood the room with water,” and the model renders the next moment of video and sound in response, continuing indefinitely until the user stops.
Runway describes the release as a direct successor to GWM Worlds, shown in December, which the company says focused on spatial consistency across long camera movements. GWM Worlds 2 adds generated audio and a new input format Runway calls WorldPrompt, which separates what stays fixed from what changes moment to moment. The fixed layer is a genesis prompt describing the scene, its subjects and the physical rules that govern them, such as gravity or camera perspective. The changing layer is a timestamped stream of text actions and camera motion, and multiple actions can overlap.
Runway says the model runs at 720p and 24 frames per second, with audio generated at 48,000 Hz, and that navigation works from either a first-person or a third-person viewpoint, with the camera and the subject being acted upon steered independently. The company also says the system started as a slower, bidirectional model finetuned on WorldPrompt data, then was post-trained into the faster autoregressive version that ships as GWM Worlds 2. All of these figures come from Runway’s own research post. The company has not published independent benchmark results or third-party evaluation of latency, frame quality, or how the model holds up over long sessions.
This is the third world model AI Insiders has covered in a week, each from a different angle. Runway’s own Solaris, an interface world model that renders app screens frame by frame, and World Labs’ Atlas both went out Sept. 2. Solaris targets interfaces rather than physical environments, so the three releases are working the same real-time video generation problem from different market entry points: interfaces, static spatial worlds, and now a combined audio-video world with no preset duration.
The durability claim is what separates this release from a demo reel. Runway says a GWM Worlds 2 session has no preset length, so the model keeps generating for as long as a user keeps issuing actions rather than producing a fixed clip on request. Paired with synchronized audio and independently controllable camera and subject, that is closer to something a game studio or simulation team could build a product around than a research showcase meant to impress and then end.
The cost of that durability is the part Runway’s post does not address. Generating continuous 720p video and 48,000 Hz audio autoregressively, frame by frame, for an unbounded session is computationally expensive, and the company has not disclosed what a minute of interactive world generation costs to run, what hardware it requires, or how long a session lasts before coherence degrades. Those numbers, not the demo clips, will decide whether GWM Worlds 2 ships as a product or stays a research preview.
Runway has not said whether GWM Worlds 2 is available to test beyond the examples in its own post, or when a public release might follow. Teams evaluating world models for game prototyping, robotics simulation or embodied-agent training should treat this release as a capability signal rather than a benchmarked product, and wait for Runway to disclose session cost and length before budgeting against it.
Runway published these findings in its own research post, “Introducing GWM Worlds 2,” on September 3, 2026.