Skip to content
Conference Open access

EVA-Bench: A Scenario-Based Benchmark for Evaluating Domain Knowledge and Agentic Capability of Foundation Models in xEVA Operations

Jul 2026 · 55th International Conference on Environmental Systems · 0 citations

Abstract

Future exploration EVA operations, especially Mars surface EVAs and some higher-tempo Artemis scenarios, will require greater crew autonomy than the International Space Station paradigm because of communication latency, limited bandwidth, and increased operational complexity. Although large language models (LLMs) and agentic AI systems show promise as onboard decision-support tools, their suitability for safety-critical EVA operations remains unclear due to the lack of domain-specific evaluation frameworks. This paper presents EVA-Bench, a benchmark designed to evaluate foundation model capabilities for exploration EVA (xEVA) support under operationally grounded and safety-relevant conditions. EVA-Bench comprises 651 tasks across six EVA scenario families, three difficulty tiers, and two complementary tracks: Single-Query (SQ) tasks for knowledge retrieval, procedural reasoning, and evidence attribution, and End-to-End (E2E) tasks for multi-step planning, tool use, replanning, and protocol compliance in dynamic mission scenarios. Tasks are grounded in a curated corpus of 84 NASA documents spanning Apollo, ISS, Artemis, EVA standards, and mishap investigations. To jointly assess capability and operational safety, the benchmark integrates a safety-sentinel framework informed by Systems-Theoretic Process Analysis, where critical protocol violations zero the final score regardless of task quality. We evaluate nine models from OpenAI, Google, and Anthropic. Results show that models perform strongly on procedural knowledge retrieval, with top SQ scores above 0.93, but degrade substantially on E2E agentic execution, where multi-step planning and contingency handling remain challenging. Results also show that smaller or mid-tier models can outperform larger models on this domain-specific benchmark, suggesting that targeted training and reasoning design may matter more than model scale alone for safety-critical EVA support. Several models with strong objective performance also trigger safety-critical violations, underscoring that raw capability alone is insufficient for operational deployment. By isolating foundation-model capability from higher-level agentic workflow design, EVA-Bench helps identify which models are most suitable for later integration into EVA decision-support systems.

Read PDF