BranchVis is presented, a prototype that represents prompts and generated images as nodes in a tree-based visualization, enabling direct navigation, unambiguous revisitation, integrated comparison, and management of complex histories.
Abstract
Generative image editing systems allow users to iteratively modify images through natural language prompts, yet most interfaces present this process as a linear conversational history. This obscures the branching nature of exploratory workflows, making it difficult for users to understand how images evolve, revisit intermediate states, and compare alternatives. In this work, we investigate how interactive visualizations of editing histories can address these limitations. We present BranchVis, a prototype that represents prompts and generated images as nodes in a tree-based visualization, enabling direct navigation, unambiguous revisitation, integrated comparison, and management of complex histories. A within-subject study (N = 12) comparing BranchVis to a conversational interface shows significantly higher usability, lower cognitive workload, and improved support for exploration, navigation, comparison, control, and understandability. These findings highlight editing history visualization as a promising direction for human-AI interaction.
Interactive visual interfaces have become an important means of controlling generative image models, enabling users to manipulate generation through prompts, direct manipulation, and a range of interactions. However, existing techniques are typically presented as independent systems, making it difficult to understand h...
Susie S. Y. Li, Mingwei Li, Remco Chang· 0 citations
Authoring and refining presentation slides is time-consuming in academic and professional settings. Although generative AI lowers the barrier to creating initial drafts, its black-box, one-way workflow often limits fine-grained control. A formative study with 10 frequent presentation authors identified trial-and-error...
Yu Fu, Yong-Qi Kang, Yu-Jia Zhou et al.· 0 citations
A striking pattern emerges: across all methods and modalities, the AI operates strictly as a spatial command executor, with collaborative and model-initiated categories entirely vacant.
Recent generative video editing models enable video content modification (e.g., changing a character) but target short clips. Extending them to full multi-shot videos requires tedious work to locate relevant content across shots, segment it into clips, craft context-aware editing prompts for each clip, and repeatedly a...
Boyu Li, Yu-Qian Zhou, Duo-Tun Wang et al.· 0 citations
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...
Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al.· 0 citations
Compositing multiple visualizations into a coherent whole remains challenging due to the vast design space and the need to balance the coverage of task-relevant data insights (e.g., trends and outliers), perceptual clarity, and aesthetic quality. In this paper, we present VisPuzzle, a task-aware method that formulates...
Zheng Wang, Zhiyang Shen, Lingyun Yu et al.· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 1, 2026
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.