Skip to content

Confidence vs. Competence: Misalignment in Judgment and Performance for Agentic Software Repair

Jul 2026 · ACM Transactions on Software Engineering and Methodology · 0 citations · 58 references

TL;DR

A persistent gap is revealed between how current repair agents assess issue conditions and how they perform during repair, highlighting the need for judgment-aware and correction-aware agent design.

Abstract

While Large Language Model (LLM) agents are increasingly applied to automated software repair, misalignment remains in how humans and agents judge issue conditions and how agents’ pre-execution self-assessments relate to repair competence. Human developers rely on diagnostic cues such as reproduction steps and stack traces to judge whether an issue is sufficiently specified, whereas LLM agents often fail to recognize missing information. We present the first systematic empirical study of misalignment between human and agent judgments and between agent judgments and repair performance. Using SWE-bench, controlled ablation experiments establish a causal link between removing human-valued cues and reduced LLM repair success. Specifically, LLM judges show limited agreement with human problem-specification ratings, and agents’ pre-execution self-assessments only weakly track repair degradation when key cues are removed. Our trajectory analysis further reveals distinct behavioral responses to missing information, while post-execution self-judgment signals add useful discriminative information when combined with behavioral traces. These findings reveal a persistent gap between how current repair agents assess issue conditions and how they perform during repair, highlighting the need for judgment-aware and correction-aware agent design.

View source

Similar papers

Preprint Aug 2026

Demystifying Agent Skills: Why They Work-Until They Don't

This work designs a contrastive study that combines controlled quantitative experiments with paired trajectory analysis and consolidates observations into a taxonomy of three high-level categories and twelve skill-use modes, showing that skills work when noisy trajectories become procedural anchors that stabilize execu...

Zhi-Yuan Jiang, Fan Huang, Hanwen Xing et al. · 4 citations · ⚡1
Preprint Sep 2026

ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers'roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies...

M. Ishfaq, Glaucia Melo · 0 citations
#artificial intelligence Preprint Sep 2026

Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance

Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns sh...

Chen Shen, Estevam Hruschka · 0 citations
#artificial intelligence Preprint Sep 2026

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the...

Hai-Qing Li, Xin-Yu Ma, Yin-Hao Wu et al. · 0 citations
Preprint Aug 2026

Grounded Checklist Partial Credit for Agent Skill Trajectories

Grounded Checklist Partial Credit (GCPC) is introduced, a human-governed and LLM-instantiated partial-credit evaluation of agent trajectories that better discriminates official PASS and FAIL outcomes than holistic judging on the shared subset.

Su-Liu Qin, Lu Yin, Xi-Lu Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.