Skip to content
Open access

Scalable and Seamless: Accelerated Virtual-Tiling for High-Fidelity Hybrid CNN–Transformer Codecs

2026 · IEEE Access · Vol 14, pp. 118263-118287 · 0 citations · 41 references

Abstract

We propose a high-capacity, end-to-end framework for large-scale image compression that addresses the trade-off between tiling scalability and perceptual quality, a challenge stemming from the patch-based processing required for high-resolution inputs, which often introduces disruptive stitching artifacts. To mitigate this issue, we present a unified framework built on three complementary components: 1) Accelerated Virtual-Tiling, which simulates boundary interactions during training to improve spatial consistency without incurring the memory cost of multi-patch encoding; 2) Seam-Targeted Distance-Masked Self-Attention, a latent bottleneck mechanism that enables information exchange across patch boundaries; and 3) Boundary-Aware Regularization, which enforces consistency at tile interfaces through an explicit loss formulation. By explicitly modeling cross-boundary dependencies, the proposed method effectively suppresses stitching artifacts while maintaining scalability to high-resolution inputs. Extensive experiments on the Kodak, JPEG AI, and CLIC 2025 datasets demonstrate competitive or superior rate-distortion performance, achieving high structural fidelity with MS-SSIM values of approximately 0.998 at high compression ratios. These results indicate that the proposed framework provides an effective solution with substantially reduced boundary discontinuities for advanced neural image compression systems based on hybrid CNN-Transformer architectures.

Read PDF