Signal Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Summary
Marionette is a world model for interactive games with articulated characters that separates state prediction, rendering and appearance synthesis instead of autoregressing pixels directly. A two-stage autoregressive dynamics model first predicts an interpretable 276-dimensional 3D world state made up of multi-entity skeletons, root trajectories and rotations. A parameter-free graphics bridge then converts that state into pose-control videos, and a control-conditioned video-diffusion model synthesizes the final RGB frames. When a mismatched action stream is forced into the system, root-aligned joint error shifts by 31 percent across 48 held-out segments, confirming that the predicted state is directly controllable. Without constraints, two generated characters drift apart to 21.2 meters, compared with about 5 meters in recorded sessions, and ground penetration appears in a third of frames, but adding a terrain collider and a separation cap to the state cuts penetration by 66 percent without changing the observation model. Routing appearance through the predicted state produces no detectable loss of visual fidelity.
Classification
Evidence 1
- Marionette: Predicting World States, Rendering Geometry, Painting Appearance arXiv (cs.AI) 2026-08-14 accessed 2026-08-20T05:08:17+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-ce3d0d260799
