Introducing our latest research: Wan-Streamer — an end-to-end omni-modal understanding and generation model designed for real-time duplex interaction.
One model. Listens, watches, understands, and responds with synchronized speech and video. Just a single native-streaming Transformer that runs a full duplex video conversation at ~550ms total latency.
Because generation is not tied to a fixed avatar rig, the model can embody anything describable in natural language — human characters, pets, anime figures, and more — all in real-time conversation.
New in v0.2: higher resolution (640×368, 25 FPS), ~200ms model-side latency.
显示更多