🕺 A Plug-and-Play 2D Motion Interface
Captioning real monocular video through 2D keypoints only — no 3D pose estimation anywhere in the pipeline.
The overlay shows the ViTPose 2D keypoints that are actually fed to the model. A_real is the 0.3M-parameter real-video adapter; the pretrained motion-language model is untouched.