Skip to content

Author

Trong-Tai Dam Vu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Noun Presence Loss with Dynamic Weighting for Compositional Text-to-Image Generation

Compositional text-to-image generation requires faithful depiction of multiple objects with distinct visual attributes. While recent inference-time optimization methods have substantially improved attribute binding, entity neglect - where specified objects are absent from the generated image - remains an unresolved failure mode. We identify that existing syntax-guided attention objectives optimize for relative modifier-noun alignment but impose no constraint on the absolute activation magnitude of entity tokens, leaving neglected entities uncorrected even when binding objectives are satisfied. To address this, we propose a noun presence loss, a hinge-based objective that penalizes entity tokens with insufficient peak cross-attention activation during inference. A loss-ratio dynamic weight scales the presence loss proportionally to the binding loss magnitude at each timestep, preventing gradient conflict while maintaining stable optimization. Our method is training-free, requires no external components, and integrates with syntax-guided cross-attention optimization frameworks. Experiments on CC-500 and T2I-CompBench across color, shape, and texture categories demonstrate state-of-the-art performance on all metrics and attribute categories, confirming that entity presence optimization consistently improves compositional faithfulness and perceptual quality across diverse prompt types.

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-Image (T2I) sub-block within the MM-DiT joint self-attention naturally encodes highly discriminative spatial localization signals. In this paper, we propose MaskFlow, a training-free framework for spatially localized image editing in rectified flow models. By strategically extracting and aggregating these attention maps from edit-relevant tokens during the standard ODE forward passes, MaskFlow automatically constructs a soft spatial mask. This mask confines semantic edits to the target region while perfectly preserving the original background. MaskFlow operates as a lightweight, plug-and-play extension without retraining the baseline model. Experiments on the FlowEdit benchmark show that MaskFlow reduces LPIPS by 28.1% relative to the baseline while maintaining competitive semantic alignment.

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.