Skip to content
Open access

LLM-Assisted UI Element Grounding and Design-Rationale Cards for Mobile Interfaces

Jul 2026 · Journal of Science Innovation and Technology Research · 0 citations

Abstract

User-interface screenshots contain dense visual and semantic structure: text labels, icons, buttons, lists, cards, toolbars, form inputs, and navigation regions frequently coexist on one screen. This study presents a compact, evidence-fused pipeline for LLM-assisted UI element grounding and design-rationale cards. The pipeline proposes UI elements from screenshot-aligned hierarchy candidates, classifies their control types, grounds natural referring expressions to element boxes, infers layout-level design tags, and produces concise rationale cards whose language is constrained by structured evidence. Experiments used 1,460 mobile UI screenshots and corresponding JSON view hierarchies, yielding 31,968 parsed nodes and 30,101 labeled elements across 25 component categories. All tasks used screen-disjoint training, validation, and test partitions. The evidence-fused detector reached 0.950 AP50 and 0.998 Recall@100 on the test screens. The evidence-fused control classifier reached 0.932 accuracy and 0.942 weighted F1. Across 2,161 generated referring trials, multi-evidence grounding reached 0.876 Top-1 accuracy, 0.979 Top-3 accuracy, and 0.928 mean reciprocal rank. Predicted component labels substantially improved layout tagging over geometry-only rules for forms, navigation, cards, image-heavy screens, and advertisements. Across 643 cards per output condition, the evidence-card generator preserved role, position, and action evidence with a contradiction rate of 0.000 under deterministic consistency checks. This last result establishes internal agreement with the evidence record, not human-rated usefulness or trust. Overall, the findings show that screenshot geometry and hierarchy metadata can be combined into an inspectable grounding-and-explanation layer, while also indicating that hierarchy dependence, rare-class imbalance, generated grounding phrases, and the absence of user evaluation limit broader deployment claims.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.