4Director Brings Rigid 3D Scene Control to Video World Models
A research system from Stability AI and UIUC uses complete object meshes and camera trajectories to steer video generation. The paper is public; implementation code is still coming soon.

Researchers at Stability AI and the University of Illinois at Urbana-Champaign have introduced 4Director, a system for directing generated video through an editable 3D scene. The preprint was submitted to arXiv on 1 October 2026. Its premise is concrete: establish the geometry and intended movement first, then let a video generator fill in the appearance.
The public repository presents demonstrations of object control, camera control and new object insertion. It currently labels implementation code as “Coming soon.” That makes this a research announcement to examine, rather than an available production tool to recommend.
Original explanatory diagram by AnnotateIt. This schematic complements the paper-derived hero; it is not an author demo or benchmark result.
What 4Director changes
Moving an object across an image is different from specifying how it should move through space. A 2D path leaves depth and orientation ambiguous. 4Director instead represents each selected object as a complete mesh and moves it as a rigid body. Object and camera trajectories share a coordinate system, making the intended scene explicit before frames are generated.
The word “world model” deserves context here. The authors use it for a generator conditioned on editable scene state; this is not a claim that the system provides an interactive simulator. For someone evaluating the work, the useful question is how faithfully a generated clip follows that state.
How it works
The pipeline starts from an image, a text prompt and prescribed trajectories. A reconstructed scene is rendered into depth frames, which condition a pretrained video generator through a Motion Adapter. Geometry supplies the scaffold; the generator supplies appearance and dynamics that the scaffold omits.
The original illustration above is a conceptual reading aid. It is deliberately not a reproduction of a paper figure or a generated demo frame. Read the methodology for the reconstruction and training details rather than treating the drawing as an implementation specification.
What can be controlled
The project demonstrates three directions for control:
- Move and rotate an object as a whole along a prescribed trajectory.
- Change the camera trajectory while keeping it in the scene’s shared coordinate system.
- Insert an object reconstructed from a separate reference image and prescribe its movement.
These are useful distinctions when watching the authors’ demonstrations. A camera orbit tests viewpoint changes; object rotation tests whether orientation follows the control; insertion tests a new scene composition. Attractive output alone would not establish that all three controls are accurate.
What the evaluation reports
The paper introduces RealCOD-Rigid, containing 20,774 clips, and Identity-Gated IoU, which evaluates placement while accounting for whether the object remains recognizable. On the authors’ 100-clip evaluation, the reported IG-IoU is 60.4 for 4Director and 54.8 for VerseCrafter. These are author-reported results from a particular protocol, not independent validation or a universal ranking.
For readers comparing systems, that distinction matters. A benchmark answers a defined question under chosen inputs and metrics. It does not by itself settle editing speed, deployment cost or suitability for a different visual domain. Those need separate evidence once an implementation is available.
Why this is interesting
The editorial signal is the separation between an explicit scene plan and the synthesis of its visual realization. It gives viewers something inspectable to compare against generated output. That could make discussions of controllability more precise than judging a plausible-looking video in isolation.
It also suggests a useful evaluation habit: inspect the commanded movement, the geometry that encodes it and the resulting frames separately. This is our interpretation of the approach, not an additional benchmark claim. For computer vision practitioners, that separation is a reason to follow the research even before downloadable tooling arrives.
What is available today
As checked on 2 October, the paper, project demonstrations and repository README are public. The repository does not yet provide an implementation release or model-weight download. Public visibility should therefore not be read as a released open-source system.
The paper also identifies a central limitation: rigid control does not prescribe articulated movement such as individual limbs. The generator decides finer dynamics. Anyone expecting a detailed character animation rig should read that limitation before extrapolating from the demos.
Follow the primary sources below for release updates. The hero adapts Figure 3 of the paper under CC BY 4.0, including an author-provided generated-video frame; it is not an independent evaluation. The diagram inside this article is our original explanation.
Sources
Media: Adapted from Figure 3 of “4Director: Controlling Video World Models with Rigid 3D Geometry” by Wei Cao et al., CC BY 4.0. Cropped and rearranged by Vision Radar; labels added. No author endorsement implied.
Source status checked on 2 October 2026. Research claims remain author-reported.