THE POTENTIAL OF AI AGENTS IN THE FIELD OF RELIABILITY AND QUALITY
Abstract
Background. Automated generation of formal specifications remains a key task in formal software verification to improve its reliability, since manual code annotation is labor-intensive and error-prone. The aim of the study is to evaluate the quality of generated annotations and identify practical limitations of the approach when using small-class models. Materials and methods. The experiment used the local models Qwen2.5-coder, CodeLlama, and DeepSeek-coder; syntactic checking of annotations was performed using SpotBugs, and formal verification was performed using OpenJML together with the Z3 SMT solver. Results. On a set of 11 OpenJML examples, medium-sized models demonstrated the ability to generate syntactically correct JML annotations; however, the proportion of fully verified specifications remained low. The main limitations were identified in the area of semantic accuracy of annotations, the inability to take side effects into account, and, in some cases, format mismatches between analysis tools. Conclusions. The proposed hybrid approach demonstrates practical potential, but its implementation in real-world development will require improving the semantics of automatically generated contracts, expanding training sets, and creating automated error correction procedures integrated into the generation cycle.