Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channels. Yet improving perceptual quality in these models still typically requires fine-tuning the network or attaching additional adapters, leav...
Federico Putamorsi, Leonardo Zini, M. Cornia et al.· 0 citations
This work introduces a human-aligned evaluation framework for text-to-SVG generation, and develops two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcem...
Marco Cipriano, Leonardo Zini, Alexandra Schild et al.· 0 citations
SLS (SVG Latent Space), a Transformer-based autoencoder that learns compact dense representations of individual SVG paths, the atomic visual elements from which any SVG image can be composed, is introduced.
Leonardo Zini, Elia Frigieri, L. Baraldi· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.