Skip to content

Reward Lightning: Fast Video Generation via Homologous Preference Distillation

Jul 2026 · arXiv.org · Vol abs/2607.03960 · 0 citations · 62 references
Computer Science

TL;DR

Reward Lightning is proposed, a unified framework that aligns and accelerates a video diffusion model within a single shared representation within a single shared representation that mitigates the gradient conflicts that arise when they are optimized over disjoint representations.

Abstract

Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize the two objectives over mismatched representation spaces, where improving one objective often compromises the other. To overcome this, we propose Reward Lightning, a unified framework that aligns and accelerates a video diffusion model within a single shared representation. Its central principle is homology: both objectives are evaluated on identical latent features, which mitigates the gradient conflicts that arise when they are optimized over disjoint representations. As a foundational component, we first introduce a latent reward model (LRM) that scores videos directly in the latent space, without decoding back to the pixel space. Building on the LRM, homologous preference distillation (HPD) reuses this shared backbone to perform adversarial distillation and preference alignment jointly, yielding few-step generators that remain faithful and well aligned. Extensive experiments demonstrate that the LRM surpasses pixel-level and latent-level reward baselines by $11.0\%$ and $14.7\%$ in preference accuracy, and that Reward Lightning generates high-fidelity videos in merely $1$ to $4$ steps, improving the average VBench score by $2.1\%$ while leading in text alignment, motion quality, and visual quality. Project page: https://reward-lightning.github.io.

View source

Similar papers

Preprint Sep 2026

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applyin...

Jiu-Zhou Lin, Jun-Long Wu, Feilong Zuo et al. · 0 citations
#computer vision Preprint Aug 2026

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

This work proposes REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher that enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost...

Yuhan Li, Fan-Gao Zeng, Sicong Kang et al. · 0 citations
#computer vision Preprint Aug 2026

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Latent-OPD is proposed, which augments OPD with trajectory-level latent distillation and introduces a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers, establishing Latent-OPD as a highly effective approach to frame-efficient video reasoning.

Aoni Shen, Yongheng Zhang, Ying-Hui Li et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Direct Preference Density Alignment for Conversational Audio Equalization

Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to...

Ioannis Stylianou, S. Shepstone, Jon Francombe et al. · 0 citations
Preprint Aug 2026

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the reliability of reward signals is constrained by the quality of human...

Naixin Zhai, Weihua Cheng, De-Xu Yu et al. · 0 citations
Conference Open access Sep 2026

Beyond Pipeline Mimicry: Expert-level Aesthetic ISP via Reward-Guided Flow

Recent deep learning-based ISP methods are primarily constrained by mimicking fixed camera pipelines, consequently struggling to achieve expert-level aesthetic quality. To this end, we propose AesISP, the first expert-level aesthetic ISP framework formulated as a reward flow model. Addressing the ill-posed nature of IS...

Tong Qiao, Kepeng Xu, Gang He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.