🕺 A Plug-and-Play 2D Motion Interface

Captioning real monocular video through 2D keypoints only — no 3D pose estimation anywhere in the pipeline.

The overlay shows the ViTPose 2D keypoints that are actually fed to the model. A_real is the 0.3M-parameter real-video adapter; the pretrained motion-language model is untouched.

Clip
0:00 / 0:00

Select a clip and generate its caption.