Comprehensive evaluations on multiple real-world benchmarks suggest that SynthRGB-T delivers superior performance and enhanced visual diversity over existing approaches, and introduces a Visual Grounding Pipeline to enhance semantic alignment.
Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.
Hyesong Choi, Daeun Kim, Song Park et al.· 0 citations
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al.· 2 citations
A Low-Complexity Cross-Modal Alignment via Projection (LCAP) network is proposed, which introduces Projective Token Compression (PTC), which leverages Mish activation and adaptive average pooling to reduce feature redundancy while enhancing discriminative information, and Positional Spatial Enhancement (PSE), which exp...
Yu-Chen Sha, Lingli Wan, Ge Yang et al.· The Visual Computer· 0 citations
Unigen-AR, a framework that pairs a general-purpose multi-modal language model (MLLM) with an efficient next-scale visual auto-regressive (VAR) decoder, is proposed, establishing visual auto-regressive modeling as a compelling and efficient backbone for unified visual generation.
Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unifie...
Weiming Zhuang, Jiabo Huang, Jingtao Li et al.· arXiv.org· 0 citations
It is shown that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness, and establish cross-branch steering as a practical tool for probing multimodal representations.
Yu Wang, Sharon Li· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.