Aug 2026· 2026 2nd International Conference on Electronic Information, Computer and Aerospace Remote Sensing (EICARS)· pp. 295-298· 0 citations· 15 references
Abstract
Data preparation, including source profiling, quality diagnosis, and cleaning, remains a labor-intensive bottleneck in data-driven applications. Traditional approaches require manual rule design for each data source and lack adaptability to heterogeneous formats. This paper proposes an LLM-driven intelligent agent for adaptive data preparation. The agent leverages a large language model as its reasoning core to perform three tasks in a closed loop: (1) automatic data source profiling through sampling-based probing, (2) data quality diagnosis and cleaning script generation via LLM code synthesis with sandbox verification, and (3) strategy memory that accumulates validated experiences as structured condition-actioneffect triples for reuse. Generated scripts are tested in an isolated environment before deployment, and both successes and failures are recorded to enable continuous improvement without model fine-tuning. This paper presents the architecture design, details the core methods, and discusses the design rationale.
Comparisons against stronger model and coding-agent competitors further indicate that both domain-specific agent runtime structure and foundation-model strength matter for autonomous data analysis.
Dynamic Retrieval-based Policy Generation (DRPG) is proposed, a framework that integrates memory-based retrieval with a dynamic policy generator, leveraging historical data and environment feedback to produce task-specific policies for continual LLM improvement.
Ting-Wei Chang, Po-Chun Chen, Hen-Hsen Huang et al.· 0 citations
This paper examines Data Agents from a harness-centric perspective, introducing a taxonomy of Data Agents and associated data environments, and identifying four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository.
Hua-Chi Zhou, Yu-Jing Zhang, Jia-He Du et al.· 0 citations
This work proposes control-data flow separation, where execution-critical control is represented as typed, validated program objects, while task-relevant language remains the optimizable data flow for agent communication.
Wentao Zhang, S. Murtaza, Junaid Ahmad Bhatti et al.· 0 citations
The DataFoundry is introduced, a framework for evolving data preparators through recursive self-improvement before large-scale data production, and it is found that recursively evolved preparators produce training data with higher downstream utility than baselines.
Ce-Hao Yang, Xiao-Jun Wu, Xueyuan Lin et al.· 2 citations
Extensive experiments show that DataMaster outperforms static baselines in most settings and surpasses full-pool training in a substantial number of cases, and simplifies data curation and removes the burden of manual strategy design.
Fanqi Zhou, Qiaosheng Chen, Zixian Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.