Skip to content

Towards Streaming Referring Video Segmentation via Large Language Model

· 0 citations · 48 references

TL;DR

This paper proposes a simple but efficient MLLM-based framework StreamingRVOS, which can extend image-level segmentation to video-level via a streaming pipeline without introducing extra parameters and achieves excellent performance in referring video segmentation.

View source

Similar papers

Open access Jul 2026

Toward Reasoning-Centric Video Object Segmentation via Multi-Modal Large Language Models

Referring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video u...

Yanyan Shao, Shuting He, Gengze Zhou et al. · 0 citations
#small language model Preprint Aug 2026

Training-Free Temporal Abstraction for General Video Understanding

STITCH is presented, a training-free method that divides a video into semantically meaningful temporal chunks that are computed once per video and reused across tasks, suggesting that reusable temporal abstraction is a promising direction for general video understanding.

Etienne Casanova, S. Brodjian, Pietro Perona · 0 citations
Preprint Sep 2026

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a text query. Since the target moment occupies only a portion of the video, PRVR requires retrieval based on fine-grained understanding beyond coarse video-level matching. However, existing methods often rely on...

Hyun Seok Seong, Woojin Jun, Subeen Lee et al. · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
Preprint Sep 2026

DynEoMT: Learning Object Dynamicity from Online Segmentation Queries

Video segmentation models recognize and track objects over time, but they do not indicate whether each segmented region moves independently of the observing camera. This dynamicity attribute cannot be inferred from semantics alone and is confounded by camera ego-motion. We introduce \method, an online framework that au...

Calvin Galagain, Martyna Poreba, François Goulette · 0 citations
Preprint Sep 2026

CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downst...

Jing-Chi Jiang, Yi-Ran Ling, Ruo-Nan Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.