DeXplicit: Making Physical State Explicit
in Dexterous Manipulation World Models
Predict how the world moves in metric 3D.
Then render what the interaction will look like.
Actual DeXplicit 14B outputs · supplied future geometry
Physical prediction first.
Visual synthesis second.
DeXplicit separates the evolution of the physical world from its visual appearance. Given initial scene geometry and a proposed hand trajectory, it predicts object motion in 3D. Those trajectories guide a video model, carrying observed appearance into future views.
From hand actions to visual futures.
A structured interface connects a lightweight dynamics network to a conditional video renderer.
Observe & act
Initial image and metric scene points, object identities, hand trajectories, and calibrated cameras.
Predict in 3D
Joint-aware actions drive shared object motion, local deformation, and bounded pose refinement.
Transport appearance
Project predicted geometry and initial-image features. Confidence identifies regions to preserve or complete.
Render the future
Synthesize conditioned video, with a causal variant that uses its own previously generated history.

Interface illustration with supplied future geometry and the 1.3B renderer. Feature colors show PCA components, not RGB appearance.
Follow the motion.
Keep the object.
Selected examples from the paper’s appendix. Explore object rotation, transport, and two-hand interaction.
These videos isolate conditional rendering: the model receives supplied future object geometry and hand actions. Actual geometry forecasts are shown separately below.
Selection and playback details
Examples were selected from complete generated sequences for visible motion, object identity, and perceptual quality. They are qualitative illustrations, not an additional benchmark or a claim of unseen-source generalization. Reference and generated video use the same frames and crop. The controls play the stored 33-frame clips; viewing speed does not establish acquisition time. No generated frame has been retouched or motion-interpolated.
Predict the physical response.
Actual saved object trajectories, conditioned on the initial scene and hand motion. Watch the recorded RGB video alongside the predicted and ground-truth 3D trajectories.
Matched states 0–32 · synchronized playback at 8 states per second. RGB is recorded ground truth; both 3D panels use the same fixed orthographic view. Colored points show the selected object; the pale outline marks its initial state.
Articulation & deformation
Selected paper examples projected into the recorded scene. Change the case to inspect the same future state across methods.
How to read the geometry visualizations
The trajectory videos show saved model forecasts, not supplied future geometry. The recorded RGB clip uses the same sample and camera as the initial observation, with frames matched to saved forecast states 0–32. All three videos share playback, frame seeking, and speed controls. RGB supplies visual context and is not a generated prediction. The 3D panels are cropped from the original saved render to remove surrounding titles and whitespace; their plotted points and camera view are unchanged. Point colors in these saved renders represent initial surface appearance; a fixed orthographic camera is used for both panels. The photographed examples overlay saved point predictions using the same calibrated camera and crop. Their background is a recorded future photograph, not generated video. Error values for these examples are selected-state 3D point errors, not full-cohort averages. Local deformation and articulation can remain underpredicted.
A visual future, one block at a time.
The causal renderer conditions on its own generated history, producing eight four-frame blocks with two denoising steps per block.
Causal-DeXplicit · supplied future geometry and hand controls · generated latent history. These selected paper illustrations overlap the causal training population; held-out quantitative evaluation is separate.
More accurate object motion.
Long-horizon moved-point error across states 1–32. Both models receive identical initial scene and action inputs.
Values reproduced from the paper’s metric 3D forecasting table. Errors are weighted by valid point-time counts; these are benchmark aggregates, separate from the selected visual examples.
Make the state explicit.
Make the future visible.
Object-centric dynamics. Transported appearance. Causal visual readout.
Explore the results