Skip to content

Elastic Token Compression for Pixel-Space Diffusion Transformers

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

The Region Token Interface (\method{}) adapts a diffusion model to these tokens, with the region count drawn at random during fine-tuning so one checkpoint serves every budget.

Abstract

Natural images concentrate their detail in a small fraction of the frame, yet diffusion models spend a full token on every patch, in every layer and at every timestep. The waste is largest in pixel-space models, with no autoencoder to absorb low-level redundancy first. Probing a pretrained pixel text-to-image transformer, we find its middle-block tokens redundant wherever the image is flat. The redundancy occupies connected, content-shaped regions, and exploiting it requires tokens with the same geometry. Cutting a Hilbert ordering of the patches provides them. Consecutive positions are always image neighbours, so any contiguous run is a connected region whose size and shape follow the content, and grouping in two dimensions becomes a cut in one. Existing reductions each lose part of this. Similarity merging scatters its groups, latent bottlenecks discard position, and skipping deletes what it should summarize. We cut where the model's features change most and pool each run into one region token. Our Region Token Interface (\method{}) adapts a diffusion model to these tokens, with the region count drawn at random during fine-tuning so one checkpoint serves every budget. \method{} leads prior reduction methods at matched budgets, matches dense quality at $2.0\times$ the speed, and stays close at $2.6\times$. The code and models are open-sourced at https://eduardzamfir.github.io/rti

View source

Similar papers

Preprint Aug 2026

MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation

MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps, delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.

Mohammadreza Hami, Mohammadreza Samadi, Chao Gao et al. · 0 citations
Preprint Aug 2026

Variable-Granularity Tokenization for High-Resolution Object Detection

ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objects and staying inside compute and memory limits. We introduce VGTok, a training-free tokenizer that sets patch granularity per region from pixels, ahead of the encoder....

Khayrul Islam · 0 citations
Preprint Aug 2026

SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching

Structural Parameter-free Affinity Regularization (SPARE), a regularizer that matches the pairwise affinities of intermediate tokens to those of the clean latents across images, is proposed, a regularizer that attains the lowest FID among parameter-free regularizers in every tested setting.

Zong-Wei Hong, Jinglun Li, Shen Zhang et al. · 1 citation
Preprint Sep 2026

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.

Xing-Yang Li, Dong-Yun Zou, Shi-Ning Zhang et al. · 1 citation · ⚡1
Preprint Aug 2026

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.

Ling-Chen Sun, Rong-Yuan Wu, Xiang-Tao Kong et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.