Skip to content
Preprint

Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026

Aug 2026 · 0 citations · 45 references
Computer Science

TL;DR

An updated evaluation of the capabilities and limitations of contemporary GenAI systems on authentic introductory OOP assessments is provided and evidence that can inform the design of programming assessments, the responsible integration of GenAI tools into software engineering education, and future studies evaluating the evolution of AI-assisted programming is offered.

Abstract

Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code. However, their performance on authentic object-oriented programming (OOP) assessments remains insufficiently understood. This study evaluates five widely used GenAI systems, ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and M365 Copilot, using programming tests and examination tasks from an introductory university OOP course. The generated solutions were assessed using the same grading criteria applied to students and compared with historical student results from the same course, as well as findings from the previous year. Common errors were also analyzed to identify recurring limitations across models. All evaluated GenAI systems achieved higher scores than the average student cohort and frequently obtained full marks on longer programming tasks. Nevertheless, they occasionally produced non-compiling code and continued to struggle with advanced OOP concepts, particularly interfaces, abstract classes, and certain inheritance-related tasks. Performance was also limited on graphics-related questions involving image interpretation. Compared with the previous year, the evaluated systems demonstrated noticeable improvements across most assessments while exhibiting several recurring error patterns. The findings provide an updated evaluation of the capabilities and limitations of contemporary GenAI systems on authentic introductory OOP assessments. They also offer evidence that can inform the design of programming assessments, the responsible integration of GenAI tools into software engineering education, and future studies evaluating the evolution of AI-assisted programming.

View source

Similar papers

Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...

Danny Brahman · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...

Chandimal Adikari, Nandika Herath · 0 citations
Book Open access Aug 2026

An Analysis of LLM Performance on Introductory C++ Programming Assignments

Background and Context. Large Language Models (LLMs) have become widely accessible to students in introductory programming courses [1], yet limited research evaluates their performance on authentic assignments with pedagogical constraints such as restricted language features and course-specific conventions. Existing be...

Iris Xu, Michał Nowak, E. Shaffer · 0 citations
Review Aug 2026

Generative AI use in Statistical Research: A Literature Review and Code Generation Case Study

GenAI can function as a research tool, but not as a substitute for methodological expertise, and has potential to increase efficiency of tasks which take advantage of its search and summarization abilities, as well as basic code debugging and algorithm formation.

Natalie Morosin, A. A. Nadi, Michael P. Wallace · 0 citations
Open access Aug 2026

Comparative Evaluation of Large Language Models in Computer Programming Education

A comparative analysis of six LLMs for generating formative feedback on introductory Java programs containing predefined defects under controlled conditions reveals substantial cross-model variation, particularly in multi-defect scenarios.

Melina Najimi, Saba Yazdani, Marzieh Ahmadzadeh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.