Skip to content
Preprint

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

Jul 2026 · 0 citations · 75 references
Computer Science

Abstract

Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.

View source

Similar papers

Open access Sep 2026

ALF-Surg: Atomic language-guided few-shot surgical phase recognition

Objective: To enhance the generalization and transferability of visual-language models (VLMs) pre-trained with extensive clinical data for surgical phase recognition (SPR) and workflow analysis tasks, this paper proposes Atomic Language-Guided Few-Shot Surgical Phase Recognition (ALF-Surg), a lightweight framework desi...

Hou-Long He, Lin Mao, Cheng-Li Song · 0 citations
Conference Open access Sep 2026

Hierarchical Conditional Energy Modeling for Medical Vision–Language Pretraining

HCE-CLIP is proposed, a vision–language pretraining framework that formulates medical image–text alignment as a hierarchical label-conditional energy modeling problem and introduces a hierarchical contradiction-based metric that quantifies logical inconsistencies between fine-grained disease predictions and higher-leve...

Cheng-Sheng Mao, Yuan Luo · 0 citations
Preprint Aug 2026

Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition

Task-specific lightweight models for surgical phase recognition excel at capturing temporal dynamics but generalize poorly under domain shift. Conversely, surgical foundation models (FMs) offer superior transferability via large-scale pretraining, yet their lack of explicit temporal modeling often yields temporally inc...

Yi-Yi Zhang, Ying Zheng, Wenxin Fan et al. · 1 citation
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Preprint Sep 2026

TEDi: Temporal Memory-Enhanced and Denoising Transformer for Surgical Instrument Segmentation

Query-based segmentation methods have shown promising potential for surgical instrument segmentation and recognition, which is essential for scene understanding and downstream tasks in computer assisted surgery. However, most existing approaches predominantly rely on per-frame predictions and overlook cross-frame tempo...

Jian Yuan, Wei-Ming Mi, Tao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.