ExplorerComputer VisionComputer Vision
Research PaperResearchia:202607.24005

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Sicheng Mo

Abstract

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registe...

Submitted: July 24, 2026Subjects: Computer Vision; Computer Vision

Description / Details

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. We ground these registers with supervision signals spanning individual agent status, global state views including bird's-eye views, and scene text. We further improve the architecture with a Mixture-of-Transformers design that uses separate weights for world state modeling and visual frame modeling. Extensive experiments in two-agent Minecraft video generation show that explicit world-state modeling improves logical consistency and generation quality.


Source: arXiv:2607.21594v1 - http://arxiv.org/abs/2607.21594v1 PDF: https://arxiv.org/pdf/2607.21594v1 Original Link: http://arxiv.org/abs/2607.21594v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Jul 24, 2026
Topic:
Computer Vision
Area:
Computer Vision
Comments:
0
Bookmark