Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from mainta...
Chu-Han Zhang, Qinghongbing Xie, Zi-Yue Wang et al.· 0 citations
AdaLens is presented, an interactive system for monitoring and steering ongoing runs that combines a storyline-based representation that unifies analytical plans, execution progress, intermediate findings, and data-column involvement with steering interactions grounded in these analytical elements for directional guida...
Yangtian Liu, Yan Miao, Shuhan Liu et al.· 0 citations
SKILL-KD is proposed, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities and consistently improves frozen student agents over fixed-model adaptation baselines.
Qi-Ming Shi, Yibo Dou, J. Zhu et al.· arXiv.org· 0 citations
BIRD-History is introduced, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems'ability to ground underspecified natural language questions using historical SQL scripts, and a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, th...
Yun-Fan Zhou, Qi-Ming Shi, Yi-Zhou Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.