This work designed a tool called dBlocks with the following features: blocks to scope content, a context manager to edit context, and inline execution to verify code, and shows how human-centered design can guide the development of LLM-integrated tools.
Abstract
With LLMs, creating software tutorials now involves steering the model's output and shaping it into a coherent, accurate learning resource, yet existing LLM tools offer writers little support for this work. By analyzing interviews with technical writers ($N=17$), we identify three requirements for how they assemble and structure multiple LLM responses, curate the context the model uses, and verify the generated content. We designed a tool called dBlocks with the following features: blocks to scope content, a context manager to edit context, and inline execution to verify code. Following a human-centered design method, we iteratively refined the design through a user study ($N=5$). In a within-subjects lab study ($N=16$) comparing dBlocks with participants'preferred workflows for LLM-assisted authoring, participants reported significantly higher confidence in the tutorials they produced with dBlocks. In addition, the tool reduced friction in verification, with writers verifying code as they drafted rather than deferring or skipping it, and helped them avoid searching long chat histories by scoping their work into blocks that kept each tutorial section and its LLM conversation together. More broadly, our work offers implications for tools that scaffold human-AI collaboration in SE workflows and shows how human-centered design can guide the development of LLM-integrated tools.
A design case study of a JupyterLab addon that delivers Socratic hints instead of direct answers, and six design hypotheses for developers of constrained AI programming assistants, addressing hint escalation, selective dialogue, context granularity, vocabulary calibration, onboarding transparency, and difficulty-aware...
Alexandre De Masi, Chen Wang, Laurent Moccozet· Proceedings of the 14th Nord...· 0 citations
When interacting with large language models (LLMs), users have limited visibility and control over how their personal information is retained, particularly after incorporation into model training. Although some deployed LLMs provide options to opt-out of training, these mechanisms do not offer guarantees about the dele...
MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.
Xue-Qing Wu, Ashwin Balasubramanian, Bingxuan Li et al.· 0 citations
Coding agents can now change code for developers, who describe goals, supply context, and respond to the agent's work. Yet prompts, screen activity, and task success each tell only part of this story. We present Say, Do, Understand, an end-to-end workflow for analyzing what developers write to an agent, what they do wh...
Yun-Han Qiao, Summit Haque, Christopher D. Hundhausen· 0 citations
This work designed two scaffolded interfaces around the same negotiation scaffold: one presented a completed AI analysis, while the other supported user-directed, incremental development, which elicited a broader repertoire of analytic requests and lower subjective effort.
Zi-Lin Ma, Suzi Jazmati, Marco Chimenton et al.· 0 citations
A scaffolded programming exercise designed to support student differentiation between good and bad GenAI code suggestions based on negative expertise–that identifying why an answer is wrong is part of developing conceptual knowledge.
J. Prather, Stephen MacNeil, Andrew Luxton-Reilly et al.· International Computing Educ...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.