DEXTEROUS MANIPULATION WORLD MODELS

DeXplicit: Making Physical State Explicit
in Dexterous Manipulation World Models

THE IDEA

Physical prediction first.
Visual synthesis second.

DeXplicit separates the evolution of the physical world from its visual appearance. Given initial scene geometry and a proposed hand trajectory, it predicts object motion in 3D. Those trajectories guide a video model, carrying observed appearance into future views.

01 / METHODOLOGY

From hand actions to visual futures.

A structured interface connects a lightweight dynamics network to a conditional video renderer.

01

Observe & act

Initial image and metric scene points, object identities, hand trajectories, and calibrated cameras.

SceneHand action
02

Predict in 3D

Joint-aware actions drive shared object motion, local deformation, and bounded pose refinement.

Metric dynamics
03

Transport appearance

Project predicted geometry and initial-image features. Confidence identifies regions to preserve or complete.

GeometryAppearance
04

Render the future

Synthesize conditioned video, with a causal variant that uses its own previously generated history.

Video readout

Actual DeXplicit 14B outputs · supplied future geometry

Inside the visual interface
One HOT3D interaction · synchronized streams
Initial observation of hands holding a plate
InputInitial observation
GeometryProjected 3D control
ActionHand joint controls
AppearanceTransported features
OutputGenerated video

Interface illustration with supplied future geometry and the 1.3B renderer. Feature colors show PCA components, not RGB appearance.

02 / CONDITIONAL VIDEO SYNTHESIS

Follow the motion.
Keep the object.

Selected examples from the paper’s appendix. Explore object rotation, transport, and two-hand interaction.

03 / METRIC 3D FORECASTING

Predict the physical response.

Actual saved object trajectories, conditioned on the initial scene and hand motion. Watch the recorded RGB video alongside the predicted and ground-truth 3D trajectories.

Focus-object motion
Endpoint error
Ground-truth RGBRECORDED
DeXplicit prediction3D
Ground-truth 3DREFERENCE
00 / 32

Matched states 0–32 · synchronized playback at 8 states per second. RGB is recorded ground truth; both 3D panels use the same fixed orthographic view. Colored points show the selected object; the pale outline marks its initial state.

Articulation & deformation

Selected paper examples projected into the recorded scene. Change the case to inspect the same future state across methods.

How to read the geometry visualizations

The trajectory videos show saved model forecasts, not supplied future geometry. The recorded RGB clip uses the same sample and camera as the initial observation, with frames matched to saved forecast states 0–32. All three videos share playback, frame seeking, and speed controls. RGB supplies visual context and is not a generated prediction. The 3D panels are cropped from the original saved render to remove surrounding titles and whitespace; their plotted points and camera view are unchanged. Point colors in these saved renders represent initial surface appearance; a fixed orthographic camera is used for both panels. The photographed examples overlay saved point predictions using the same calibrated camera and crop. Their background is a recorded future photograph, not generated video. Error values for these examples are selected-state 3D point errors, not full-cohort averages. Local deformation and articulation can remain underpredicted.

04 / CAUSAL GENERATION

A visual future, one block at a time.

The causal renderer conditions on its own generated history, producing eight four-frame blocks with two denoising steps per block.

1initial observation
→
→
32future frames
2steps / block

Causal-DeXplicit · supplied future geometry and hand controls · generated latent history. These selected paper illustrations overlap the causal training population; held-out quantitative evaluation is separate.

05 / QUANTITATIVE CONTEXT

More accurate object motion.

Long-horizon moved-point error across states 1–32. Both models receive identical initial scene and action inputs.

Retrained PointWorldDeXplicitLower is better · mm

Values reproduced from the paper’s metric 3D forecasting table. Errors are weighted by valid point-time counts; these are benchmark aggregates, separate from the selected visual examples.

DEXPLICIT

Make the state explicit.
Make the future visible.

Object-centric dynamics. Transported appearance. Causal visual readout.

Explore the results