Skip to content
Open access

Graph Neural Networks with Code Structure Embeddings for Model-Agnostic Multi-Class Detection of AI-Generated Programs

Sep 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

GraphCSE, a detector for unseen generators of AI-generated code, represents a program as a heterogeneous code graph with four edge relations and fuses structural and semantic node channels through a learned per-node gate before relation-aware graph attention, restoring model-agnostic detection.

Abstract

Most detectors of AI-generated code answer a binary question: human or machine. We study the multi-class problem, naming the generator family behind a program, and the open-set problem that follows from it, in which the generator was never seen during training. Our detector, GraphCSE, represents a program as a heterogeneous code graph with four edge relations and fuses structural (syntax-type) and semantic (token) node channels through a learned per-node gate before relation-aware graph attention. On a controlled synthetic corpus of 2,500 Python programs spanning a human class built from eight style archetypes and four generator-family proxies, GraphCSE reaches 92.1 ± 0.6% macro-F1, the strongest learned model we test and 11.2 points above a GCN on identical graphs, though character n-gram TF-IDF remains slightly ahead in-distribution (93.4 ± 0.5%). The paper's central findings concern unseen generators. First, a negative one: on single programs, all five rejection rules we test (closed-set argmax, centroid and k-NN distance, energy, maximum softmax probability) fail on held-out families, with mean AUROC no better than 0.65: a tight unseen style inside the human envelope is indistinguishable from an unusual human. Second, a constructive one: judging a set of programs from the same source by a two-sided anomaly test on embedding statistics, calibrated only on human data, restores model-agnostic detection, reaching 86.4-97.3% mean detection at ten programs per source (AUROC up to 0.972) in the graph embeddings versus 53.6% in raw TF-IDF space. Provenance of unseen generators is a property of collections, not snippets.

Read PDF

Similar papers

Preprint Aug 2026

Learning Spectral Representations of Code through Latent Graph Learning for Generalizable Cross-Language Code Clone Detection

Current code clone detection (CCD) methods rely on fixed, language-specific graph representations like abstract syntax trees (ASTs) or program dependency graphs (PDGs). Because functionally identical code fragments can yield wildly different structures, these rigid graphs produce non-discriminative spectra that perform...

Mohsen Hesamolhokama, Alireza Sadeghi, Kousha Moeini et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection

Malware detection is a critical task in cybersecurity, and graph neural networks over control flow graphs have shown promising results for it. However, detectors are usually evaluated on a random split of a corpus collected over a single period, which cannot show how well a model generalizes to later samples. This stud...

M. Sajeed, Mayukh Mondal, Md. Ashraful Hossen Akash · 0 citations
#graph neural networks Preprint Aug 2026

Can Graph Learning Learn Circuits?

Graph Circuit Learning is introduced, a supervised, amortized framework that trains a GNN across multiple model--task pairs and applies it to unseen cases and preliminary results suggest that graph machine learning offers a natural and potentially powerful perspective on circuit localization.

Chester Tan, Moritz Lampert, Courtney Maynard et al. · 0 citations

CodeBERT-SENet: Adaptive Syntax-Semantic Fusion via Gated Attention for Python Bug Detection and Localization

This research proposes the first end-to-end multi-task architecture that jointly detects Python syntax errors and localizes their exact line position by adaptively fusing CodeBERT’s semantic embeddings with handcrafted syntactic features via a lightweight gating mechanism, without relying on Abstract Syntax Trees.

Rozali Ilham, Bahtiar Imran, Hasan Basri et al. · 0 citations

IMoKGNN: Dual-Stream Fusion of Generic and Task-Specific Language Model Features for Graph Neural Networks

Text-Attributed Graphs (TAGs) are prevalent in various real-world scenarios, where each node is associated with a text attribute. Representation learning on TAGs relies on a comprehensive understanding of both the textual attributes and the topological connections. Recent works have enhanced graph neural networks (GNNs...

Hao Yan, Chao-Zhuo Li, Jun Yin et al. · 0 citations
Preprint Aug 2026

GraphAlignCoder: Aligning Program and Proof Graphs for Code Generation

GraphAlignCoder is introduced, a training framework that transfers explicit correctness structure into code generation and consistently outperforms the base model, code-only SFT, and CodeRL across all benchmarks.

Yue-Ke Zhang, Zihan Fang, Kevin Leach et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.