ModiGen: A Large Language Model-Based Framework for Modelica Component Generation in Multi-Domain Systems
Abstract
Modelica is an industry-standard language for modeling and simulating complex cyber-physical systems, playing a critical role in digital twin development. Driven by the escalating demand for complex system modeling, there is a growing interest in leveraging the advanced code generation capabilities of Large Language Models (LLMs). To explore this potential, we establish a comprehensive benchmark to systematically evaluate LLMs in generating Modelica component models. Our evaluation reveals that current LLMs struggle to construct simulation-ready components due to strict acausal semantics, with failures primarily stemming from: (1) syntactic violations, (2) semantic hallucinations, and (3) physical inconsistencies. To overcome these challenges, we propose ModiGen, a framework designed to promote syntactic correctness along with structural and physical consistency in Modelica generation. ModiGen features supervised fine-tuning to facilitate the internalization of domain knowledge. Crucially, it employs Modelica-specific structural retrieval-augmented generation to inject precise topological constraints and physical laws. Furthermore, it incorporates a fully automated, feedback-driven refinement mechanism that transforms simulation diagnostics into fine-grained instructions, guiding LLMs to resolve specific syntactic and semantic errors. Experimental results demonstrate that ModiGen achieves up to 18.65% improvement in the functional validation pass rate over baselines, significantly enhancing the reliability of component generation.