Preprint
Aug 2026
Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting
Structured video prompting is introduced, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding.
Sadegh Mohammadian
· 0 citations