Jul 2026· Electronic Proceedings in Theoretical Computer Science· Vol abs/2607.21209, pp. 515-528· 0 citations· 23 references
Computer Science
TL;DR
This work proposes a framework that outlines how to produce comprehensible explanations for policy-aware agents, or agents which have rule-enforcing policies incorporated in their decision-making framework, and is designed using insights from the social sciences on how to produce good explanations.
Abstract
In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal. As these systems grow more prevalent in our day-to-day lives, there has been an increased need to add explainability features which can provide an account for an agent's behavior. We therefore propose a framework that outlines how to produce comprehensible explanations for policy-aware agents, or agents which have rule-enforcing policies incorporated in their decision-making framework. This framework is designed using insights from the social sciences on how to produce good explanations. It is implemented in the Answer Set Programming language while using Python to assist with information extraction and natural-language translation. Because these agents incur penalties when violating policies, we are able to leverage these penalties to detect undesirable events in scenarios that are counterfactual to the agents'original actions. This lends itself to creating contrastive explanations (e.g.,"the agent performed this action because, had it not, undesirable event X would have occurred."), which formulate the core component for our explainability framework. The framework is evaluated using a survey wherein human participants provide feedback on our program-generated explanations.
The evolution of artificial intelligence has enabled autonomous AI agents to reason, plan and make decisions in complex and dynamic environments. AI that behaves human-like is not a single model or a rule-based system. It is a collection of agents that collaborate to find solutions to issues that a single AI model is u...
Praveen Dommalapati· International Conference Com...· 0 citations
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing...
CO Tiffany, Wen Zhang, E. Bagdasarian et al.· 0 citations
It is argued that governing such agents is a runtime problem -- not a model-alignment problem and not a build-time problem -- and five primitives are derived from the questions that must be answered before an action takes effect and after it has: discovery, identity, governance, attestation, and supply chain.
This work introduces defeasible action rules that capture the typical outcome of action, without committing to what happens in exceptional circumstances, and defines an inference operator for defeasible action rules inspired by rational closure, adapted to the dynamic setting.
This work presents a post-hoc XAI framework that transforms a lengthy agent's execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior, outperforming naive LLM-generated explanations.
Vittoria Vineis, Fabiano Veglianti, Lorenzo Antonelli et al.· 0 citations
The results show that giving LLM agents access to world model beliefs improves task performance under partial observability, while remaining complementary to existing simulation-based world models.
Shubham Kumar, Harshit Kumar, Narendra Ahuja et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.