This work introduces XstrAI, an audience-aware multi-agent framework that treats local explanations as fixed evidence and structures how it is communicated to each target reader, and evaluates XstrAI on diabetes and stroke risk prediction against 11 baselines.
Abstract
Feature-attribution methods such as SHAP provide useful evidence about individual model predictions, but their numerical outputs are rarely sufficient for audiences with different expertise, goals, and risks of misinterpretation. In medical AI, the same local explanation must reach patients, clinicians, and data scientists through markedly different forms of communication, and naive verbalization through large language models (LLMs) is prone to weak grounding, conflation of attribution with causal language, and outputs that are persuasive without being faithful to the underlying model evidence. We introduce XstrAI, an audience-aware multi-agent framework that treats local explanations as fixed evidence and structures how it is communicated to each target reader. Each prediction case is encoded as an immutable structured representation, shared identically across audiences so the underlying evidence remains fixed. Generation is factored into three specialized LLM agents responsible for audience-aware planning, linguistic realization, and validation for grounding, attribution consistency, communicative risk, and audience appropriateness, with a bounded revision loop triggered on detected inconsistencies. We evaluate XstrAI on diabetes and stroke risk prediction against 11 baselines, ranging from direct verbalization to a re-implementation of a state-of-the-art narrator. The evaluation combines an intra-narrative regime measuring fidelity to SHAP evidence with an extra-narrative regime assessing audience appropriateness through reference corpora, multi-family LLM judges, and a survey with target readers. In both evaluations, XstrAI's narratives are consistently assigned to their intended audience by independent judges, and preferred over all baselines on Clinician and Patient audiences, with competitive performance on Data Scientist, where audience-conditioned single-prompt baselines lead.
Persuasive argument generation requires modeling audience beliefs, rhetorical strategies, and factual grounding. Despite recent advancements, existing methods remain largely audience-agnostic and fail to integrate strategy selection to improve persuasiveness. To bridge this gap, we propose Argus, an agent-based framework that operationalizes classical rhetoric for persuasive writing. At its core, a Theory-of-Mind (ToM) Reasoner constructs an explicit dual mental model of the audience's beliefs and values to guide downstream decisions. This representation conditions a component-aware planner that decomposes the argument into subtopics, assigns fine-grained rhetorical functions (logos, pathos, ethos), and triggers strategy-guided evidence retrieval at planning time. Finally, a refinement module iteratively targets and resolves multi-dimensional weaknesses without quality regression. We evaluate Argus across three diverse benchmarks using both automated pairwise Elo and LLM-as-judge metrics. Results show that Argus consistently outperforms strong baselines across multiple backbone models, achieving top rankings and the highest overall scores. Targeted simulation experiments further validate its effectiveness in shifting resistant audience stances.
Experiments across multiple long-context narrative question answering and claim verification settings show that ClueWeaver substantially improves local end-to-end language models while providing evidence coverage and paragraph-referenced reasoning traces.
Ji-Hao Zhu, Zhi-Wei Yang, Wen-Xiao Zhang et al.· 0 citations
Inferring speaker relationships from spoken conversations is an important step towards socially aware speech understanding. However, this task remains underexplored, and supervised modeling is costly to train and scale. At the same time, existing inference-time LLM approaches provide limited structure for handling subtle, distributed, and multimodal relational cues that may support multiple plausible interpretations. To address these limitations, we introduce a training-free multi-agent reasoning framework that organizes inference through structured interaction among LLM agents, allowing relationship judgments to be proposed, challenged, and adjudicated without task-specific training. We instantiate this framework with two complementary designs. We propose Multi-Role Multi-Agent Debate as a task-specific adaptation of standard multi-agent debate for speaker relationship inference, assigning agents complementary roles or social-theory-grounded perspectives rather than a single undifferentiated viewpoint. In contrast, we introduce Multi-Agent Compete, a competition-based protocol that compares agent judgments through pairwise adjudication, eliminates weaker candidates, and retains the most defensible one. We evaluate these methods on the Seamless Interaction dataset across different modality settings, covering both binary classification and fine-grained relationship-detail prediction. Results suggest that they improve over zero-shot and existing multi-agent baselines in most cases. Human evaluation further suggests that this task is challenging even for people. LLM methods can sometimes outperform human annotators in text-included settings but are less competitive in the audio setting. Together, these findings suggest that relationship inference benefits from structured inference-time interaction among agents, while acoustic cues are not yet fully captured by current models.
Yao-Han Guan, Yen-Ju Lu, Yu-Zhe Wang et al.· 0 citations
Automated fact-checking systems still fall short of producing explanations that mirror the depth and structure of expert human reasoning. In this work, we propose a multi-agent framework that integrates five specialized linguistic agents covering polarization, linguistic style, argumentation, plausibility, and contextual framing with web-based evidence retrieval, synthesized by a supervisor agent into structured reports resembling professional fact-checking outputs. We evaluate the framework on a dataset of fact-checked Brazilian news through a classification benchmark and two further quantitative studies of explanation quality, addressing: (1) Do the generated reports elicit reader confidence comparable to reports written by professional fact-checkers? and (2) Which explanatory dimensions most influence reader confidence? The classification benchmark shows the framework performs competitively with strong baselines. A blinded within-subjects study with 95 participants, analyzed via Linear Mixed Models, shows that post-verification confidence reaches levels statistically indistinguishable from expert-written reports, with plausibility and analytical depth as the strongest predictors of confidence gain and depth being especially important for implausible claims. Complementary LLM-as-a-judge experiments corroborate these findings, showing the framework’s explanations are consistently preferred for depth, persuasion, and plausibility.
Pedro Henrique de Oliveira Silva, L. Santos, L. Marinho et al.· Proceedings of the 37th ACM...· 0 citations
An evaluation method is proposed that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs and proposes a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations.
Eran Hirsch, David Wan, Han Wang et al.· 0 citations
Chimamanda Ngozi Adichie’s metafictional short story, “Jumping Monkey Hill” is used as a laboratory to explore the relationship between fictionality and believability in a situation where rhetorical cues about a text’s ontological status are misleading and how literary ways of reading might inform engagement with AI texts.
J. Freed· Narrative Inquiry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.