Over the past two years, AI video models have been competing on realism, resolution, and duration. But no matter how impressive the results look, we remain passive viewers: we press play, watch the clip, and it ends.
AlayaWorld
@alayastd is attempting something fundamentally different. Instead of generating a fixed video, it generates a world that continues to unfold as you move through it.
These three demos show the same journey toward a green village rendered in three distinct styles: photorealistic, oil painting, and line art. As the camera moves forward, the model continues generating the road, fences, trees, and distant village. This is not simply an existing video with different filters applied. The environment is generated continuously along the camera trajectory, allowing the scene to develop as the user explores it.
AlayaWorld streams video at 720p and 24 FPS while supporting camera movement and viewpoint control. The real breakthrough is not just image quality. Once generation becomes fast enough to respond within an interactive loop, the user is no longer merely watching a video. They become a participant inside the generated world.
The world can also respond to new instructions. During generation, users can introduce prompts that trigger spells, summon characters, create explosions, or transform the environment. Most video models follow an initial prompt and produce a predetermined clip. AlayaWorld can respond to changing intent while the world is still running, allowing subsequent events to evolve according to the user’s commands.
Generating an attractive frame is relatively easy. Maintaining a coherent world over time is much harder. As a video model repeatedly predicts the next frame, small errors can accumulate until roads, buildings, and objects begin to distort or disappear. AlayaWorld combines spatial memory with compressed historical context, helping the model remember both where things are and what has already happened. This enables stable generation lasting more than one minute while improving consistency when the camera leaves an area and later returns.
This may be the next step for AI video: not simply generating a longer movie, but generating a world that can be explored, changed, and interacted with.
AlayaWorld is developed by Alaya Lab. The team is progressively releasing its inference code, training code, and datasets, with an online experience expected to launch near the end of the month.
Project page: