Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clich\'es or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clich\'es into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clich\'ed. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
Hongyu Luo, Hexi Wang, Huihao Jing et al.· 0 citations
Immersive and believable NPC dialogue requires characters that feel intentional. They remember specific information about themselves, follow through on their goals, and stay true to their personalities across long conversations. We introduce Post-Thinking, a technique that maintains a rolling reflection trace across chat turns. After each response, effectively in the dead time between LLM queries, the model generates a trace reflecting its current goals, emotional state, and narrative intentions. This trace is kept in context and conditions the next response, directing the conversation while designed to add zero perceivable latency to the end user. Critically, each trace is generated with the previous n reflection traces still visible in context, allowing the character’s inner state to compound and evolve naturally throughout extended conversations. As a preliminary study, we synthetically annotated conversations from seed datasets and interviewed human experts to assess their quality in preparation for fine-tuning. We expect the reflection generation pass to help actively ground the character as it encourages the model to explicitly surface aspects of the character definition most relevant to the current moment, counteracting the prompt drift that typically degrades consistency over long exchanges.
Keegan Carey, Hexi Wang· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.