When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Avyay M. CasheekarHariganesh Tangirala
Aug 2026
Artificial Intelligence
Abstract
Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcome finality and cross-unit separation. Outcome finality requires that later events cannot change the claimed result, while cross-unit separation requires that earlier runs cannot change the relevant conditions of later ones. The endpoint establishes neither condition by itself, and the two can hold independently. Waiting for a delayed outcome may settle the label even though its state remains available to another run. Isolation may prevent carryover even though the scored outcome remains unresolved. We develop a completion argument that identifies the evidence needed for each decision. A final success or failure label is justified only when every relevant effect is resolved or bounded tightly enough to fix the outcome. Any remaining uncertainty must be reported. First, in a controlled replay with fixed agent actions, we find that endpoint and terminal labels differ for every nonzero-delay operation and that a delayed write changes the next run's score under shared state but has no such effect after namespacing or verified reset. Second, in a review of ten public protocols, we find that reset or deliberate retention is documented explicitly more often than unfinished operations or evidence for separate scoring. Finally, we propose an open-effects record for operations and resources that may remain relevant after the endpoint, their status, and their possible effects on the scored outcome or another run.
Investigating how experienced developers use agents in building software, including their motivations, strategies, task suitability, and sentiments finds that while experienced developers value agents as a productivity boost, they retain their agency in software design and implementation out of insistence on fundamental software quality attributes.
An adaptive surrogate modeling method for problems with very high-dimensional spatio-temporal outputs is developed that combines exploration and exploitation to improve the surrogate model accuracy with the fewest possible runs of the expensive physics-based model.
B. Kapusuzoglu, S. Mahadevan, Shunsaku Matsumoto et al.· Structural And Multidiscipli...· 17 citations
An adaptive jailbreak attack framework for systematic evaluation of both cascaded pipelines and end-to-end large audio-language models under a unified experimental setting that achieves consistently higher attack success rates across diverse audio-based LLM systems.
Linghan Huang, Bo Li, Huaming Chen et al.· 12 citations· ⚡2
This review provides a systematic literature review of LLM-based Verilog code generation, analyzing 102 papers (70 published and 32 high-quality preprints) from SE, AI, and EDA venues and outlines a roadmap highlighting potential opportunities in LLM-assisted hardware design.
This work introduces Behavior-Outcome Freedom (F), a pre-synthesis diagnostic of signed behavior-outcome rank mismatch, and formalizes its candidate-conditional role through Signed Anchor-Rank Transfer, which preserves validated capability resources, removes runtime orchestration, and conditionally inherits pipeline guidance using a calibrated rule over F.
Binyan Xu, Dong Fang, Haitao Li et al.· arXiv.org· 10 citations
Simulation results confirm the effectiveness and benefits of DMs in generating neighbor velocity estimates in a four-UAV swarm coordination task using Deep Reinforcement Learning (DRL), and explore the integration of DMs with RL and DT.