Skip to content

A Typography Benchmark for Co-Creative Graphic-Design Agents

· 0 citations · 21 references

TL;DR

A typography-focused benchmark of 12 tasks grounded in a layered-composition corpus of 989 real-world design templates and 2,568 text elements is introduced to quantify how well future design partners can perceive, reproduce, and manipulate the ty-pographic layer designers care about most.

View source

Similar papers

Preprint Aug 2026

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clich\'es or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clich\'es into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clich\'ed. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.

Hongyu Luo, Hexi Wang, Huihao Jing et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated from real-world creative imagery, organized into three levels and ten categories, with each sample paired with two prompts of differing specificity. For evaluation, we develop a hybrid framework within an MLLM-as-judge paradigm, combining a multiple-choice question (MCQ) based protocol of 9,594 questions across four levels of metaphorical fidelity with a dimension-based scoring protocol along three perceptual dimensions. Extensive evaluation of 11 representative T2I models reveals that even the strongest proprietary models struggle with compositional structuring and cross-domain mapping, key aspects of metaphorical expression, highlighting visual metaphor generation as an important frontier for future T2I research.

Chuer Chen, Zi-Chen Wang, Yi He et al. · 0 citations
#computer vision Preprint Sep 2026

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain''for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator''to synthesize standalone visual assets. Operating under an ``imagine first, then act''closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.

Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al. · 0 citations
Open access Sep 2026

A Scalable Art-Direction Methodology for Live-Service Visual Production

Live-service games depend on a continuous supply of visual content, including characters, environments, user interfaces, animations, visual effects, seasonal events, and marketing materials. When a studio grows to dozens or hundreds of artists who support several products simultaneously, art direction that relies on the constant personal involvement of a single strong specialist loses stability. Quality becomes uneven, key people carry an excessive load, and errors surface at late, expensive stages. This article describes a practitioner-derived methodology that treats art direction as a reproducible production system. The system links visual strategy with organizational architecture, capacity planning, technical standards, multilevel quality control, specialist development, and LiveOps. The proposal rests on two distinctive elements. The first element scales artistic vision through a distributed decision-making structure in which project and discipline leads hold defined authority within set boundaries. The second element plans visual production across the full life cycle of an asset, from brief and scope estimation to integration, validation, release, and reuse. A comparison with published approaches positions the methodology against the Art Bible tradition, Scrum, industry conference practices, and production-tracking software. Boundaries of applicability and directions for empirical validation are outlined.

Dmytrii Yeremenko · 0 citations
Conference Jul 2026

Sketchify: An AI-Powered Visual Website Builder using Large Language Models and Canvas-based Wireframe Interpretation

Building a functional website traditionally demands proficiency in HTML, CSS, and JavaScript, or costly subscriptions to no-code platforms. Sketchify addresses this barrier by providing a canvas-based wireframe drawing interface that converts hand-drawn sketches directly into production-ready HTML/CSS websites using Large Language Models (LLMs). A custom spatial layout interpreter processes drawn elements and classifies them into semantic web components such as navbars, hero sections, cards, and footers. The enriched layout descriptor is forwarded to the Groq API (Llama 4 Scout) to generate single-file HTML output styled in one of six design paradigms: Glassmorphism, Skeuomorphism, Neo Brutalism, Claymorphism, Minimalism, and Liquid Glass. The system further supports multi-page generation, a natural-language post-generation chat editor, a Supabase-backed project dashboard, voice-command input, and a template gallery. Acceptance testing confirmed successful website generation within five minutes of onboarding. Generation latency is under fifteen seconds at near-zero infrastructure cost. Benchmarking against Pix2Code and Sketch2Code confirms superior output quality and design flexibility without task-specific model training.

H. Babu, Densy John Vadakkan · 0 citations
Review Aug 2026

Trade-offs in Data Color Palette Design Tools

Designing a color palette for data requires designers to balance multiple constraints, including accessibility and aesthetics. Color palette tools support this process through features including direct manipulation, automated palette generation and evaluation, previews, and so on. Despite their prominence, relatively little is known about how these different mechanisms shape design across contexts. We conducted an exploratory think-aloud crowd work study with 40 self-identified designers. Each participant used one of four palette tools selected to span different interaction modalities to complete a series of accessibility- and aesthetics-oriented design tasks. We observed two preliminary patterns. First, tool differences were more pronounced in accessibility-constrained tasks. Second, even when accessibility was not explicitly required, some tools produced more accessibility-friendly palettes and prompted more accessibility-oriented thinking. In this tool genre, then, system design shapes outcomes both via built-in functionality, as well as by directing designers'attention toward particular constraints and design considerations.

Shi-Yi He, A. Mcnutt · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.