Skip to content

StableFlow: Real-Time 4K Video Super-Resolution With Robust Feature Propagation

Sep 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 9988-10000 · 0 citations · 48 references
Medicine

Abstract

Real-time 4K video super-resolution (VSR) requires the effective reuse of temporal information under strict latency constraints, typically relying on temporal alignment and real-time reconstruction. However, imperfect alignment introduces perturbations into the temporal recursion, which can accumulate over time and degrade reconstruction quality—an effect widely observed in recurrent VSR pipelines but rarely analyzed explicitly from a dynamical perspective under real-time constraints. In this work, we formulate real-time VSR inference as a recurrent dynamical process, embedding alignment within the state update. This formulation enables explicit analysis of how alignment-induced perturbations are introduced and propagated during inference. Motivated by this analysis, we propose StableFlow, a stability-guided real-time VSR framework. Specifically, StableFlow introduces an Alignment Gain-aware Alignment Module (AGAM) to enhance the utility of temporally aligned features, a State-aware Perturbation Control Filter (SPCF) to suppress unreliable propagated information, and a Temporal Propagation Control (TPC) loss to regulate long-term recurrent state evolution. Our method combines efficient temporal alignment with state-aware perturbation control and a temporal propagation control loss to stabilize long-term recurrent behavior. Experiments demonstrate that StableFlow achieves real-time 4K performance (over 40 FPS on an NVIDIA RTX 3090) for <inline-formula> <tex-math notation="LaTeX">$4\times $ </tex-math></inline-formula>upscaling from <inline-formula> <tex-math notation="LaTeX">$960\times 540$ </tex-math></inline-formula> inputs to <inline-formula> <tex-math notation="LaTeX">$3840\times 2160$ </tex-math></inline-formula> outputs, with only 321K parameters and 77.45G FLOPs, while maintaining competitive reconstruction quality.

View source

Similar papers

Aug 2026

Event-Guided Online Video Super-Resolution

Event-guided video super-resolution (VSR) leverages high-temporal-resolution event streams to address motion blur, rapid dynamics, and poor illumination that challenge frame-only VSR methods. However, most existing approaches emphasize reconstruction quality while overlooking real-time performance and computational eff...

Ze-Yu Xiao, Xinchao Wang · 0 citations
#artificial intelligence Preprint Aug 2026

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Disti...

Yu-She Cao, Shikun Feng, Ru-Xiang Duan et al. · 0 citations
Preprint Sep 2026

DecoGS: Adaptive Static-Dynamic Decoupling of 3D Gaussians for Free-Viewpoint Video Streaming

Streaming 3D reconstruction demands both speed and temporal fidelity, goals that existing methods undermine by updating every Gaussian every frame, even in static regions. We present DecoGS, a method for efficient online training of 3D Gaussians from streaming videos. Unlike prior methods that update the entire scene i...

Idil Sulo, Alexey Supikov, Ilke Demir et al. · 0 citations
Preprint Sep 2026

LVMT: Video Mask Transformer for Long-term Video Segmentation

Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their in...

Narges Norouzi, Niccolò Cavagnero, Idil Esen Zulfikar et al. · 0 citations
Preprint Sep 2026

Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement

Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to t...

Zi-Yin Huang, Sik-Ho Tsang, Xin Qin et al. · 0 citations
Preprint Aug 2026

Following Motion for Sequential Modeling in Video Frame Interpolation

State Space Models (SSMs) have surfaced as a promising architecture in Video Frame Interpolation (VFI), as they can capture long-range dependencies with linear computational complexity. However, their predefined scanning order limits their effectiveness in modeling the dynamic motion trajectories inherent in VFI proble...

Jaehyun Park, Nam Ik Cho · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.