Skip to content

ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation

Dec 2025 · arXiv.org · Vol abs/2512.07328 · 5 citations · 49 references
Computer Science

TL;DR

Experiments demonstrate that ContextAnyone improves both identity and fine-grained appearance consistency over existing reference-conditioned baselines while maintaining motion characteristics close to the underlying text-to-video generator.

Abstract

Text-to-video generation has advanced rapidly, yet preserving a character's holistic appearance from a single reference image remains challenging, particularly when the character undergoes large pose, motion, and scene changes. Existing reference-conditioned approaches primarily treat the reference image as a conditioning signal, which can weaken fine-grained appearance information as reference and noisy video tokens interact during denoising. We propose \textbf{ContextAnyone}, a context-aware diffusion framework that instead treats the reference as an explicitly preserved appearance anchor. Our key idea is to jointly reconstruct the reference image and generate the target video within a shared diffusion transformer, providing direct supervision for preserving identity and fine-grained appearance throughout denoising. To maintain the reference as a stable source of appearance information, we further introduce asymmetric information flow that allows video tokens to selectively access reference tokens while preventing noisy video features from propagating back to the reference branch. We complement this design with Gap-RoPE, which separates the positional representations of the reference and generated video tokens. Experiments on a benchmark constructed from OpenVid-HD demonstrate that ContextAnyone improves both identity and fine-grained appearance consistency over existing reference-conditioned baselines while maintaining motion characteristics close to the underlying text-to-video generator.

View source

Similar papers

Open access Sep 2026

Bridging Text and Motion: Generative AI Models for Video Synthesis

Text to video generation has advanced significantly in recent years, largely due to the development of extremely sophisticated diffusion models. In this work, we present a novel ap- proach to producing excellent video content based on descriptions by utilizing diffusion tech- niques. Using a multi-stage diffusion proce...

Mohammad Shahnawaz Shaikh · 0 citations
Preprint Sep 2026

The Price of Consistency: Exploiting Visual Anchors for Multimodal Jailbreaking in Video Generation

The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significa...

Peng Li, Qian-Qian Xu, Yangbangyan Jiang et al. · 0 citations
Preprint Sep 2026

TexTailor: Texture-Preserving Video Virtual Try-On via Adaptive Garment Conditioning

Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to...

Zi-Jing Qin, Jun Zhou, Rui-Cheng Zhang et al. · 0 citations
Preprint Sep 2026

MSR: Multiple Subject Reference for Video Generation

Conditioning a video generator on multiple images requires preserving appearance while associating each reference with its intended role. We present MSR (Multiple Subject Reference), a slot-aware conditioning scheme for LTX-based video generation. Each reference image is independently encoded as a static clip and repre...

Guan-Nan Li, Jia-Ji Chen, Jing-Yuan Liao et al. · 0 citations
#large language models Preprint Sep 2026

ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits

ViTAL-X, a lightweight model that equips frozen image-text backbones with temporal awareness while preserving their foundational spatial knowledge, achieves state-of-the-art performance and demonstrates that targeted, high-quality temporal alignment provides a highly efficient alternative to pure scaling.

V. SethuramanT., Savya Khosla, O. Susladkar et al. · 0 citations
Preprint Aug 2026

TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models

A lightweight Text Encoder Alignment framework that fine-tunes only the text encoder while keeping the generative backbone fully frozen, and achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts.

Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.