Skip to content
Preprint

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

A warping-error bound is derived that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows, and ChordVideo achieves competitive temporal consistency and source preservation using 10--60 model steps per clip.

Abstract

One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce \textbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by \textbf{78\%} and flicker by \textbf{49\%}, improves CLIP frame consistency by \textbf{9--10 points}, and increases background PSNR by about \textbf{1.5,dB}, while retaining \textbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using \textbf{10--60$\times$ fewer model steps per clip

View source

Similar papers

Jul 2026

OSVE: One Step Video Editing with One Step Diffusion Models

OSVE is presented, the first framework to successfully adapt one-step Text-to-Image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency.

Habin Lim, Gyeong-Moon Park · 0 citations
Preprint Aug 2026

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically d...

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al. · 0 citations
Book Open access Aug 2026

VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent Anchors

An uncertainty-aware variational latent anchoring module that dynamically selects informative frames and compresses cross-frame latents into a compact set of semantic anchors that achieves state-of-the-art inversion fidelity, editing quality, and temporal consistency, while reducing memory and runtime compared with pri...

Zhangkai Wu, Xuhui Fan, Zhongyuan Xie et al. · 0 citations
Jul 2026

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

This paper proposes ElasticTTT, a novel framework that preserves the prior generative distribution and rescues generative elasticity in standard TTT, achieving state-of-the-art performance on one-shot video editing.

Yueyi Liu, Chi Zhang, Sen Cui et al. · 1 citation
Preprint Aug 2026

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.

Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al. · 5 citations
Preprint Aug 2026

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint...

Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.