Skip to content
Open access

An Empirical Study on the Effectiveness of Large Language Models for Automated Software Development Tasks

2026 · Economic Insights: Trends and Challenges · 0 citations

Abstract

This research investigates the possibility of replacing junior and intermediate programmers with artificial intelligence (AI)-based tools in the code generation process. To inspect the capabilities of these tools, the Qwen Coder tool was used to conduct the tests. The methodology used in the study is based on 10 tests, each of which presents a certain peculiarity in relation to the level of a junior or intermediate programmer. The tests include scripts, desktop applications, and web applications. Additionally, these tests also considered requirements for innovation, database handling, image processing, specific software infrastructure, and, finally, the integration of third-party services. The performance evaluation of the applications generated by the Qwen Coder tool was based on criteria that target the functional correctness of the code, the degree of understanding of the requirement, response time, the complexity of the generated application, the number of files, and the programmer's level of satisfaction. The results obtained showed that the tool generates functional code in a fraction of the time compared to a programmer's abilities, having the capacity to understand simple or medium-level tasks. In most cases, the generation time was under 5 minutes, and the applications were executed without major interventions from the programmer. A well-defined aspect in the paper addresses the necessity for the programmer to understand the specific workflow of the tool with which they generate code. In the case of the Qwen Coder tool, it requires the integration of an auxiliary tool called GitHub for managing the repositories where the code is deposited by Qwen Coder. A second aspect highlighted in the paper refers to prompt engineering, as the quality of the generated code is directly dependent on the formulation of the prompt and the level of detail within it. The results of the study show that these AI tools successfully replace entry-level and intermediate programmers. This research investigates the possibility of replacing junior and intermediate programmers with artificial intelligence (AI)-based tools in the code generation process. To inspect the capabilities of these tools, the Qwen Coder tool was used to conduct the tests. The methodology used in the study is based on 10 tests, each of which presents a certain peculiarity in relation to the level of a junior or intermediate programmer. The tests include scripts, desktop applications, and web applications. Additionally, these tests also considered requirements for innovation, database handling, image processing, specific software infrastructure, and, finally, the integration of third-party services. The performance evaluation of the applications generated by the Qwen Coder tool was based on criteria that target the functional correctness of the code, the degree of understanding of the requirement, response time, the complexity of the generated application, the number of files, and the programmer's level of satisfaction. The results obtained showed that the tool generates functional code in a fraction of the time compared to a programmer's abilities, having the capacity to understand simple or medium-level tasks. In most cases, the generation time was under 5 minutes, and the applications were executed without major interventions from the programmer. A well-defined aspect in the paper addresses the necessity for the programmer to understand the specific workflow of the tool with which they generate code. In the case of the Qwen Coder tool, it requires the integration of an auxiliary tool called GitHub for managing the repositories where the code is deposited by Qwen Coder. A second aspect highlighted in the paper refers to prompt engineering, as the quality of the generated code is directly dependent on the formulation of the prompt and the level of detail within it. The results of the study show that these AI tools successfully replace entry-level and intermediate programmers

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.