CivBench is presented, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP), and two interface-level metrics are introduced that the environment makes measurable: Proactive Monitoring Rate (PMR) and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns.
Abstract
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30 to 75 turns, and in 7 of 20 detectable defeats they failed to query within the 20 turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48.2% and 65.8% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://github.com/lmwilki/civ6-mcp
This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent...
Yingpeng Ma, Jianhao Yan, Bei-Ning Shi et al.· 1 citation
This survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile, and provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.
Yi Bin, Xiao-Yang Yuan, Hao Zeng et al.· 1 citation
Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models to command decision-making in RTS games, enabling agents to read textual game states an...
Xin-He Tian, Xiao-Yue Zhang, Zi-You Zhang et al.· 0 citations
SimLife is introduced, a scalable platform for simulating long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues with audio that evaluates long-context pattern understanding: the ability to infer latent behavioral rules from weeks or months of everyday observations.
Run Peng, Zinnia Nie, Jing Ding et al.· 0 citations
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs...
Tianyou Wang, Chong-Yang Gao, Ke-Zhen Chen et al.· 1 citation
WorldBench is presented: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions, and Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and ot...
Leonardo Ranaldi, Sherrie Shen, Jushi Kai et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.