Skip to content

Category

artificial intelligence

6,270 papers

WebXSkill: Skill Learning for Autonomous Web Agents

This work introduces WebXSkill, a framework that bridges a grounding gap with executable skills, each pairing a parameterized action program with step-level natural-language guidance, and finds that better skill deployment mode depends on a model's plan and execution capability.

Zhaoyang Wang, Qianhui Wu, Xuchao Zhang et al. · 17 citations

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

A trajectory-aware evolutionary search method, T-MAP, which leverages execution trajectories to guide the discovery of adversarial prompts and enables the automatic generation of attacks that not only bypass safety guardrails but also reliably realize harmful objectives through actual tool interactions.

Hyomin Lee, Sangwoo Park, Yumin Choi et al. · 8 citations · ⚡1

Retrieval-Augmented LLM Agents: Learning to Learn from Experience

Simple episodic retrieval is established as a strong foundation for agent memory and retrieval-aware fine-tuning as a practical and effective framework for building agents that learn to learn from experience.

Thomas Palmeira Ferraz, Romain Deffayet, Vassilina Nikoulina et al. · 6 citations
#artificial intelligence Preprint Jan 2026

ShardMemo: Scope-Before-Routing for Agentic Memory Retrieval

SHARD-MEMO is presented, an agentic memory system built on scope-before-routing: metadata predicates first identify the admissible shards, and a learned router then selects a small number of them for shard-local approximate nearest neighbor retrieval, separating hard admissibility from learned relevance ranking.

Yang Zhao, Chengxiao Dai, Mengyi Kou et al. · 0 citations

Error-Driven Scene Editing for 3D Grounding in Large Language Models

This work introduces DEER-3D, an error-driven framework that diagnoses grounding failures and generates targeted counterfactual training supervision via a structured "Decompose, Diagnose, Edit, and Retrain"loop, and underscores the effectiveness of targeted, error-driven scene editing in bridging linguistic reasoning with spatial grounding in 3D LLMs.

Yue Zhang, Zun Wang, Han Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2025

Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models

Two test-time monitoring strategies are proposed: reasoning expression monitoring and hidden states monitoring, that reduce token usage by 62.7-93.6%, substantially improving efficiency and reliability while largely preserving accuracy.

Qingjie Zhang, Yujia Fu, Yang Wang et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2025

Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion

Persuasion Duality is identified: reasoning enhances an agent's persuasive power while simultaneously increasing its resistance to persuasion, and a prompt-level adversarial argument detection method is proposed that consistently improves agent robustness.

Haodong Zhao, Jidong Li, Zhaomin Wu et al. · 4 citations · ⚡1

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

D2Snap, an algorithm to downsample the DOM, premised on preserving actionability and actionability-discriminating features, is proposed and evaluated using a snapshot-variant web agent (GPT-4o) on a dataset sampled from Online-Mind2Web.

Thassilo M. Schiepanski, Nicholas Piël · 4 citations
#artificial intelligence Preprint Open access Sep 2026

When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, and encourage clients to continue expressing themselves. However, existing counseling dialogue systems and evaluation frameworks often favor explicit, content-rich replies, overlooking the interactional value of brief counselor utterances. This paper presents a systematic cross-lingual analysis of minimal responses across multiple counseling dialogue datasets. We develop a two-stage filtering method based on utterance length and content, followed by contextual verification using a large language model (LLM). Our analysis shows that minimal responses are common in human-collected datasets but substantially underrepresented in LLM-generated ones. We further evaluate current LLMs in manually curated dialogue contexts where human counselors used minimal responses. The results show that strong commercial LLMs are capable of generating minimal responses when explicitly instructed, but still struggle to determine when such responses are appropriate. Counseling-specific models trained on synthetic data perform particularly poorly, tending instead to produce longer and more information-rich responses. Moreover, LLM-based response-quality evaluation may undervalue minimal responses, even when they are interactionally appropriate.

Zhiyang Qi · 0 citations
#artificial intelligence Preprint Aug 2026

LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space

Inspired by the Big Five model's dimensional view of personality, a framework that reframes authorial writing characteristics as coordinates within a unified and interpretable space is proposed, which improves authorial expressiveness while preserving semantic fidelity.

Jinghui Zhang, Lang Gao, Ao Li et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.