Skip to content

Author

Sahar Vahdati

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Open access Aug 2026

Games as Experimental Philosophy: Human and Artificial Learning in ARC-AGI-3 Worlds

How do humans and artificial agents learn to act in a world whose rules they do not yet know? We use LS20, a novel ARC-AGI-3 game, as an artificial world: a microcosm whose mechanics must be discovered through exploration and improvisation alone. We operationalise a methodological loop: observe how humans explore, adapt, and become flexible; extract principles of adaptive coupling; build artificial systems embodying those principles; and let the comparison reveal where the artificial model still diverges. By analysing the completed trajectories of 18 human players, we identify a descriptive bottleneck level for each player: a level that consumes a disproportionate share of total actions, after which many trajectories become more efficient. The clearest temporal effect is that players pause 2.1× longer after actions that change a distal reference pattern, consistent with state checking after a causal intervention. Our current artificial agent also solves LS20, but requires about 2,400 actions across 17 attempts, roughly 3.6× the mean human action count (human range 405–1111 actions). We interpret the gap as a difference in how interaction is organised under uncertainty and resource pressure: human players appear to turn visual differences into action-testable regularities, while the agent still relies on slower explicit probing. We frame these patterns through the protocognition taxonomy (Rodriguez-Vergara and Husbands, 2026) and the bodily mindedness framework (Parvizi-Wayne and Montefiore, 2026), arguing that what emerges is neither mindless flow nor reflective deliberation, but a form of skilful, flexible engagement whose computational underpinnings remain an open empirical question. Data/Code available at: ARC-AGI-3 games: https://three.arcprize.org; project data, agent traces, scripts, and figures are available from the authors upon reasonable request.

Andrei C. Aioanei, Syeda Khushbakht Batool, Alexander Aston et al. · 0 citations
Preprint Jul 2026

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal while maintaining action efficiency because every move counts. We formalize these environments as parameterized rendered deterministic Moore machines and introduce Tycho, a coding-agent system that constructs and uses game-specific models during interaction. Tycho separates actionable observations from intermediate animation, level-completion, and game-over frames. From this structured history, an agent can model, test, plan with, repair, or bypass a free-form executable hypothesis. In one matched public-set run per policy, we compare four orchestration policies on all 25 public games using Claude Opus 4.8 under matched inference budgets. Actor-requested delegation to a model builder obtains the highest observed mean Relative Human Action Efficiency (RHAE), 88.49. With this selected policy, GPT-5.6 Sol and Opus 5 both reach 100.00 RHAE and complete all 183 levels. Their game-balanced first-run human-replay midranks are 98.5 and 100.0. Opus 5 uses 61% fewer scored actions than the aggregate official human baselines. Automatic repair after verification failures produces models that reproduce observed transitions much more accurately, yet reaches only 83.07 RHAE. Transition match indicates whether a simulator reproduces observed dynamics, not whether it has identified the objective or improves the next action. Strong play also requires deciding when to construct, repair, use, or bypass a model. We call this joint problem active abstraction: generating a testable model from costly interaction and deciding when acquiring or using it is worth its cost.

Jens Lehmann, Andrei C. Aioanei, Sahar Vahdati · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.