Skip to content

Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

Jul 2026 · arXiv.org · Vol abs/2607.09892 · 0 citations · 44 references
Computer Science Engineering

TL;DR

DenseAR is extended to a unified model that handles multiple modalities and imaging tasks within a single backbone that unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods.

Abstract

We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid with progressively denser strides naturally captures the transition from global structure to fine detail. This addresses two limitations of existing autoregressive models at once: the slow inference of raster-order autoregression, which DenseAR avoids by predicting multiple tokens in parallel, and the heavy cost of multi-scale approaches, which need long, multi-resolution token sequences to achieve coarse-to-fine prediction. Building on our efficient framework and the flexibility of autoregressive modeling, we further extend DenseAR to a unified model that handles multiple modalities and imaging tasks within a single backbone. We validate DenseAR on both medical and natural images. On multi-contrast brain MRI, a single DenseAR model unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods. On ImageNet, DenseAR improves class-conditional generation quality (FID and IS) over both a single-grid baseline without stride ordering and a multi-scale tokenizer-based baseline.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless...

Meimingwei Li, Stefan Andreas Baumann, Felix Krause et al. · 0 citations
Aug 2026

SCALAR++: Efficient Controllable Generation via Scale-wise Visual Autoregressive Learning

This work proposes a Scale-wise Conditional Decoding mechanism, which projects semantic signals from a frozen vision encoder into scale-specific layers of the VAR backbone, and introduces a Unified Control Alignment strategy (SCALAR-Uni) to handle diverse control modalities within a single projection space.

Ryan Xu, Dong-Yang Jin, Shawn Chen et al. · 1 citation
Preprint Aug 2026

VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting

Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like human skeletons is still unexplored. VARPose is proposed to adaptively densify 2D sparse poses, thereby enriching the anatomical information available for 3D lifting mode...

Kai Pu, Tiantian Yang, Danel Zeng · 0 citations
Preprint Aug 2026

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

MTAR introduces multi-token prediction (MTP) to alleviate the sparsity and myopia of traditional NTP by imposing joint supervision on multiple future tokens; employs token-level contrastive regularization (TCR) to explicitly enhance the separability of sampled token representations and thereby improve representation di...

Guo Niu, Xiongfei Yao, Teng Wang et al. · 0 citations
Preprint Aug 2026

XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling

High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a challenging process that depends heavily on teacher-model quality. In this pape...

Jin-Xiu Liu, Xuan Liu, Kang-Fu Mei et al. · 0 citations
Jul 2026

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

The key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling and replacing global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the...

Dehao Hao, Kaiyi Zhang, Tanghui Jia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.