Skip to content

Interactive Generative Motion Editing via Scheduled Inpainting

Jul 2026 · Computer graphics forum (Print) · Vol 45 · 0 citations · 49 references
Computer Science

TL;DR

This work introduces scheduled inpainting, a method that enables interactive generative motion editing, a novel paradigm unifying motion synthesis and editing by leveraging generative models, and extensively validate the approach by comparing with four baselines, conducting ablations of the design, and reporting user feedback.

Abstract

Motion editing is central to VFX and game development, where it is used extensively to modify and augment existing movements to conform to new environments or changes in artistic direction. While traditional motion editing can do small modifications, it cannot accommodate larger structural edits, resulting in visual warping artifacts that require authoring new motion. Conversely, recent advances in large-scale generative modeling have unlocked newfound capabilities for authoring entire movements by directly manipulating sparse spatial constraints. While impressive at creating new movements, these methods lack the capability to preserve and edit existing motion interactively. In this work, we introduce scheduled inpainting, a method that enables interactive generative motion editing, a novel paradigm unifying motion synthesis and editing by leveraging generative models. Scheduled inpainting is a simple yet powerful inference-based technique that enables fine-grained spatiotemporal control over the balance between preserving the original motion and generating new content. By building atop generative models that support direct manipulation, our system allows artists to interactively refine existing animations while ensuring results remain natural and consistent with the learned motion distribution. Scheduled inpainting is versatile and supports many editing applications, such as extending, stitching, and compositing different clips. Finally, we extensively validate our approach by comparing with four baselines, conducting ablations of our design, and reporting user feedback.

View source

Similar papers

Preprint Aug 2026

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible...

Yu-Qian Zhou, Zhenghong Zhou, Zongze Wu et al. · 0 citations
Preprint Aug 2026

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically d...

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al. · 0 citations
Conference Open access 2025

Enhanced Text-to-Image Editing with Multi-Step Control and Explainability

: Text-driven image editing has advanced significantly in generating and modifying visual content. Existing approaches often face challenges in maintaining visual coherence across sequential edits and providing informative rationales for alterations. This approach develops an improved text-to-image editing system that...

S. R, S. K., S. Harish et al. · 0 citations
Jul 2026

ViP-Rig: Visual-Prompted Controllable Rigging

ViP-Rig is a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones into a frozen pretrained autoregressive generator.

Zihan Qin, Ming-Ze Sun, Yifan Mao et al. · 1 citation
Preprint Aug 2026

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

Wan-Animate-2 is presented, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer and achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely.

Guangyuan Wang, Liucheng Hu, Dechao Meng et al. · 2 citations
2025

Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling

In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancements have led to notab...

Chenghao Xu, Guangtao Lyu, Jiexi Yan et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.