Comparative Analysis of Rewriting-Based Large Language Model-Generated Synthetic Code Detection
Abstract
The case of synthetic code detection (i.e., written by a human or generated by AI) becomes more important in the field of education and security. Current code detection tools that depend on probabilistic and classification-based approaches have limitations in detecting various programming styles. To address this challenge, this study proposes an extension of a previous work on a zero-shot rewriting similarity-based detection approach by utilizing the encoder-decoder architecture of CodeT5. Additionally, the study uses a problem–answer dataset written in Python to evaluate the proposed approach against several code variants. The experimental results show that all four CodeT5 configurations achieved a modestly higher average AUROC (0.5805) than the GraphCodeBERT-based configurations evaluated in this study (0.5731), though the margin is small and was not tested for statistical significance. Detection performance also varied substantially across code variants, especially in cases of structural modifications that reduce detection effectiveness. The results prove that code written by AI and code written by humans have distinct, detectable rewriting patterns. This confirms that the rewriting-based approach works well, even under highly demanding stylistic tests.