Skip to content

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Jul 2026 · arXiv.org · Vol abs/2607.14681 · 2 citations · 62 references
Computer Science

TL;DR

This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships.

Abstract

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordinate information from multiple visual sources accurately. We identify a critical deficiency in existing approaches. Existing editing instructions lack explicit reference relationships, and most multimodal large language models (MLLMs) cannot generate them reliably. To address this problem, we propose ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing. Our key insight is embedding reference tokens at semantic positions to eliminate ambiguity and establish precise bindings between visual attributes and their sources. We develop ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships. We further develop ReBind-Edit, which enables lightweight adaptation of text-to-video models to coordinate multiple references by binding visual attributes to their designated sources. Extensive experiments demonstrate that ReBind substantially outperforms general-purpose MLLMs in instruction quality and achieves state-of-the-art performance among open-source methods on reference image conditioned video editing. Our project webpage: https://rebind-mrv2v.github.io/.

View source

Similar papers

Preprint Aug 2026

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples, is introduced and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency.

Bojia Zi, Xiaoyan Yang, Yu Zhou et al. · 0 citations
Preprint Aug 2026

VicEdit: Learning to Edit Videos from Visual In-Context Examples

This work proposes Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair, and curates VicEdit-400K, the first large-scale dataset for visual in-context video editing.

Yu-Ji Wang, Teng Hu, Yuheng Chen et al. · 0 citations
Conference Jul 2026

Dual-path attention modulation for training-free text-guided image editing

A training-free Dual-path Attention Modulation (DAR) framework that decouples semantic edits while preserving source image structure is proposed and Adaptive Self-Attention (ASA) and Adaptive Cross-Attention (ACA) modules that dynamically regulate attention replacement are introduced.

Tong Cui, Jie Yang, Kai-Ru Li et al. · 0 citations
Aug 2026

PHA-Net: Prototype-based hierarchical alignment network for text-video retrieval

A new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities and introduces multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment.

Xiaolun Jing, Kezhao Yin, Xinxing Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.