Skip to content
Conference

HierDiff: hierarchical saliency-guided diffusion for template synthesis

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 143511Y - 143511Y-13 · 0 citations · 47 references
Engineering

Abstract

Automated template generation plays a vital role in digital advertising, filmmaking, and virtual environments, yet challenges remain in spatial control particularly in maintaining overall coherence between salient foreground and background regions. We propose HierDiff, a novel saliency-guided diffusion framework that unifies spatial structure and semantic consistency to enable controllable template generation without requiring costly fine-tuning of the underlying UNet. HierDiff introduces three lightweight modules: (1) a Saliency-Guided Spatial Attention Module, which utilizes saliency maps to direct attention toward important regions and enhance spatial awareness; (2) a Hierarchical Semantic Encoder, which captures multi-scale semantic cues from saliency inputs to provide rich layout context; and (3) an Adaptive Fusion Module, which dynamically injects spatial and semantic priors into UNet features through feature modulation. These components collectively enhance structural coherence, maintain generative diversity, and provide fine-grained layout control. Extensive experiments on a dataset with saliency annotations demonstrate that HierDiff significantly improves generation quality, yielding a substantial reduction in FID compared to baseline methods, while also enhancing perceptual similarity and other evaluation metrics. The proposed framework advances controllable and spatially-aware template synthesis, showing strong potential for structure-aware and layout-guided content generation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.