Advancing Machine-generated Text Detection: A Comprehensive Evaluation of Transformer-based Models
Abstract
—Improved fluency in large language models has intensified the need for accurate detection of machine-generated text. This study evaluates transformer-based models using an improved version of the Conference on Computational Linguistics 2025 (COLING 2025), Generative Artificial Intelligence (GenAI) Content Detection Task 1 dataset, which was carefully preprocessed to enhance label quality and balance. All models were trained under a unified protocol to ensure fair comparison and robust evaluation. Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%. These results highlight the effectiveness of advanced transformer architectures for distinguishing human-written and machine-generated text. Despite these gains, performance degradation under domain shift and highly paraphrased inputs remains a challenge.