Pufibara is presented, an agent harness that maintains persistent engineering state across revisions, associates execution and simulation evidence with the candidate that produced it, and makes submission an explicit agent action to evaluate end-to-end Modelica agent workflows.
Abstract
AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents different requirements from general-purpose code generation in software engineering, because correctness depends not only on syntax and executability but also on physical consistency and scenario-dependent behavior. We study this challenge in Modelica, an equation-based modeling language in which a model may compile and simulate while still violating its intended physics or engineering requirements. Across successive revisions, an agent may lose track of requirements or rely on simulation evidence produced by an outdated candidate. To address this challenge, we present Pufibara, an agent harness that maintains persistent engineering state across revisions, associates execution and simulation evidence with the candidate that produced it, and makes submission an explicit agent action. To evaluate end-to-end Modelica agent workflows, we also propose a source-grounded method for constructing realistic and independently evaluable tasks. We use this method to build the 232-task Modelica Agent Workflow Benchmark, spanning Model Repair, Model Generation, and Model Tuning. Each submitted candidate is scored by a benchmark-owned evaluator outside the agent loop. We compare Pufibara with Claude Code as complete harnesses under two matched large language model (LLM) backends. With DeepSeek v4 Flash, Pufibara passes 202 tasks, compared with 185 for Claude Code. With Claude Sonnet 5, Pufibara passes 202 tasks, compared with 187 for Claude Code. Under the repository-reported token accounting, Pufibara records 76.4%-82.5% lower logical-token totals. Its sequential runtime is 6.1%-58.4% lower. These findings show that, even under matched LLM backends, complete agent harnesses can differ substantially in both task success and resource use for physical system modeling.
This work deploys AutoMOOSE as a agentic software, complementing the prior work which focused on development of the agentic tool, and describes its software framework and architecture through Use Case and logical views of the 1+5 architectural-views model.
Sukriti Manna, Henry Chan, Subramanian K. R. S. Sankaranarayanan· 0 citations
It is argued harness engineering is a distinct, effective discipline for agent reliability, but its cost is model-dependent and must be measured, not assumed.
Traditional application architectures assume behavioral logic authored in advance, leaving reachable behavior largely bounded by explicit code and workflow rules. Large language model (LLM)-based agents challenge this assumption by enabling runtime reasoning and autonomous action to become part of application behavior,...
David Luo, Alberto Leon-Garcia· IEEE Access· 0 citations
This work argues that physical execution should be a first-class modeling concern and proposes physical execution plans for Model-Driven Engineering, which complements meta-models with execution plans that declare layouts, relations, indexes, access paths, and target runtimes without changing model meaning or modeling...
Francisco Martínez-Lasaca, J. de Lara· Proceedings of the ACM/IEEE...· 0 citations
This work asks what a CAE simulation agent still needs beyond a generic harness, and finds that with information access and repair budget held fixed, a single-agent harness matches or beats multi-agent specialized systems.
The more typical feature of agentic AI systems is dynamic, multistep workflows where autonomous components plan, reason, and communicate with external tools and data sources in a series of iterations. Such flexibility increases capability but also brings nondeterminism which is inherent and where the same inputs can re...
Ankur Gupta, Karan Gupta, Divyakumar Deepak Savla et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.