Skip to content
Preprint

Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing

Aug 2026 · 1 citation · 53 references
Computer Science

TL;DR

RefineCut is introduced, which, unlike workflow systems that wrap a prompted frontier model, trains a compact open-weight planner for it, which matches or exceeds its frontier teachers.

Abstract

Practical video editing is not only pixel generation: an editor must turn a brief, a clip pool, music metadata, and hard constraints into an executable timeline. We study this decision layer as \emph{executable video-editing planning} and introduce RefineCut, which, unlike workflow systems that wrap a prompted frontier model, trains a compact open-weight planner for it. The planner edits a typed timeline through structured patches covering clip selection, trimming, ordering, transitions, and duration and music alignment; a deterministic verifier applies each patch and checks it against an explicit constraint ledger. Because editing has no single ground-truth repair, we do not imitate teachers directly: RefineCut replays every multi-teacher branch through the verifier and keeps verifier-best repairs as supervision. A second stage, RefineCut-Evo, lets the student score its own repairs with the verifier and a task rubric and trains on high-margin preference pairs, so the final $8$B planner runs in a closed verifier loop with no teacher calls at inference. On RefineCut-Bench ($3{,}578$ tasks, $7{,}971$ captioned clips, $499$ music tracks, explicit ledgers), verifier-replayed distillation lifts the planner from $0.620$ to $0.858$ on the protocol-specific Video-Editing Score and RefineCut-Evo reaches $0.924$; the gain transfers to Llama-3.1-8B and GLM-4-9B, and in the same closed loop the $8$B planner matches or exceeds its frontier teachers. Code and RefineCut-Bench are publicly released; see the Data Availability statement.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from select...

Gunin Gupta, Nirmit Arora, Pavan Tankala · 0 citations
Preprint Sep 2026

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constr...

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar et al. · 0 citations
Review Sep 2026

Grounding with Confidence: Controllable Generative Video Temporal Grounding

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate...

Jin-Hao Chen, Ben-Lei Cui, Rui-Jian Jia et al. · 0 citations
Preprint Aug 2026

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-image rewards is inefficient because a poor edit does not reveal whether additional optim...

Hao-Xiang Cao, Jiajiong Cao, Xuan-Pu Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting...

Gautam Rajendrakumar Gare, Si-Ying Li, He-Wei Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.