The RenderMatte dataset is constructed, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets that features exact strand-level alpha annotations and diverse background composites, demonstrating a scalable path toward high-fidelity matting in open-world scenes.
Abstract
Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editability of downstream creation workflows. However, precise alpha estimation in open-world scenes remains challenging because real foregrounds exhibit highly diverse appearances and opacity patterns. This makes existing methods struggle with semantic ambiguity and fine-grained opacity variation, especially in sparse boundary regions that are fragile and difficult to supervise. To address this gap, we present RenderMatte, a trimap-guided matting framework that adapts FLUX.1 Kontext through full-parameter fine-tuning, leveraging image editing priors for structure-preserving alpha prediction. During supervised adaptation, an alpha-edge objective preserves the latent flow-matching signal while strengthening pixel-space boundary supervision. We further introduce group-relative alpha alignment for post-training. It compares multiple mattes sampled under the same trimap condition using matting-specific rewards for alpha accuracy, boundary fidelity, trimap compliance, and compositional consistency. To overcome the lack of precise edge annotations, we construct the RenderMatte dataset, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets. It features exact strand-level alpha annotations and diverse background composites. Experiments show state-of-the-art performance across all benchmarks, demonstrating a scalable path toward high-fidelity matting in open-world scenes.
This work proposes a region-based feature enhancement framework built upon a Topology-aware Segment Graph (TSG) that achieves superior color accuracy and temporal stability compared to state-of-the-art frameworks, particularly in scenarios involving complex character motion and topological variation.
Bin Huang, Haoran Mo, Chengying Gao· IEEE Transactions on Visuali...· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
The proposed MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions, incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it.
Rui Xu, Yang Yong, Shun-Zi Yang et al.· 0 citations
IAB edited achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics and shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions.
Chiranjeev Chiranjeev, Muskan Dosi, M. Vatsa et al.· 0 citations
PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.
Yu-Feng Chi, Hui-Min Ma, Fan Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.