Skip to content
Preprint

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages, is introduced, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.

Abstract

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously. To address these challenges, we introduce TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages. Our framework decouples semantic generation from pixel rendering: a vision-language model (VLM) handles visual understanding, cross-lingual translation, and structured visual generation, while a diffusion model performs background inpainting and pixel-level refinement, followed by deterministic rendering to synthesize the final image. Based on this formulation, we develop a three-stage post-training framework, where supervised fine-tuning (SFT) establishes the image-to-code mapping, privilege-gap weighted self-distillation (PWSD) improves the learning of style and layout tokens, and reinforcement learning with verifiable rewards (RLVR) further optimizes task-level performance. We further introduce TransAnyDataset and TransAnyBench, a multilingual dataset and benchmark for e-commerce image translation. Extensive experiments demonstrate competitive performance against cascaded pipelines, open-source end-to-end models, and closed-source image editing systems, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.

View source

Similar papers

Preprint Sep 2026

T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation

Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In applications such as product labeling, signage, and interface design, target text should be rendered within a designated text-bearing region w...

Yan Wang, Xin-Yi Hou, Wei-Guo Lin et al. · 0 citations
Jul 2026

InnoText: A Unified Model for Visual Text Generation and Editing

This work proposes InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model, and introduces a Font Size-Aware Modulation module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a...

Hao-Wei Liu, Runze He, Jian Lu et al. · 0 citations
Preprint Aug 2026

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations, generalize to out-of-domain benchmarks, confirming that effective VTC hinges on ali...

Tianyu Liang, Xiangxi Zheng, Yilin Wang et al. · 2 citations
#computer vision Preprint Sep 2026

On the Design Fundamentals of Pixel Text Representation Learning

This work investigates the fundamental design principles required for robust visual text representation learning and trains Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples.

Chaohao Yuan, Rui-Feng Yuan, Zhuoxu Huang et al. · 1 citation
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Jul 2026

TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

This paper introduces the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images, and proposes TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts.

Meng Xie, Li Zeng, Hang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.