Skip to content
Preprint

A Scalable Vector Graphics Latent Space

Aug 2026 · 1 citation · 46 references
Computer Science

TL;DR

SLS (SVG Latent Space), a Transformer-based autoencoder that learns compact dense representations of individual SVG paths, the atomic visual elements from which any SVG image can be composed, is introduced.

Abstract

Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a continuous, dense, and invertible latent space for vector representations, the kind of foundational building block that Variational Autoencoders and their descendants have long provided for raster images. We introduce SLS (SVG Latent Space), a Transformer-based autoencoder that learns compact dense representations of individual SVG paths, the atomic visual elements from which any SVG image can be composed. By modeling SVG commands, coordinate data, and visual properties within a unified BPE-based token vocabulary, SLS learns fixed-size latent representations that jointly capture structure and appearance, and can be decoded back into valid, style-consistent SVG paths with high fidelity. The resulting embedding space is robust, invertible, and structured: embeddings lie on a unit hypersphere, enabling efficient similarity search, composition, and downstream conditioning through simple vector-space operations. Finally, we demonstrate that SLS generalizes across diverse tasks reducing their FLOPs by over 150 times compared to token-based approaches, and establishing a general-purpose latent foundation for vector graphics research.

View source

Similar papers

Preprint Aug 2026

V-RAE: Rethinking Video Latent Spaces for Generation

V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, and tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality are introduced.

Ming-Hui Guo, Sheng-Qiong Wu, Hao Fei · 3 citations
#artificial intelligence Preprint Aug 2026

Composed Historical Image Retrieval by Modeling Temporal Representations

This paper proposes Temporally Decomposable Image Representations (TDIR), a representation learning algorithm that decomposes historical photographs into separate date and content components through orthogonal subspaces, and defines and proves the conditions under which such a decomposition is achievable.

Adrià Molina Rodríguez, Oriol Ramos Terrades, J. L. Canet · 0 citations
Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level ma...

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
Preprint Aug 2026

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

AURORA-LM is introduced, a continuous-latent diffusion language model that separates the construction of a decodable text representation from the modeling of its distribution, and achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText free generation and XSum su...

Jiajun Liang, Yu-Ling Liao, Yu-Kang Cao et al. · 2 citations
Preprint Sep 2026

Generative Late-Interaction Embeddings For Visual Document Retrieval

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retrai...

M. Eltahir, Talal Aloushan, Rose Khairoalsendi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.