Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition
This work introduces \the authors', which compresses frame-to-frame changes into latent actions and predicts next-frame features during end-to-end video--language alignment and achieves competitive recognition with faster inference and smaller INT4 accuracy drops than V-JEPA2/2.1.