Skip to content
Open access

Selecting and Combining Large Language Models in Scalable Code Clone Detection

Aug 2026 · ACM Transactions on Software Engineering and Methodology · 1 citation · 80 references

TL;DR

Findings indicate that ensembling approaches can be statistically significant and effective on larger datasets, where the best-performing ensemble improved performance by 37% over its individual LLMs on the commercial large-scale code.

Abstract

Source code clones pose risks ranging from intellectual property violations to unintended vulnerabilities. Effective and efficient scalable clone detection, especially for diverged clones, remains challenging. Large language models (LLMs) have recently been applied to clone detection tasks. However, the rapid emergence of LLMs raises questions about optimal model selection and potential LLM-ensemble efficacy. This paper addresses the first question by identifying 76 LLMs and filtering them down to suitable candidates for large-scale clone detection via an LLM-encoder framework known as SSCD. The candidates were evaluated on two public, industry-defined datasets, BigCloneBench, and a commercial, large-scale dataset. No uniformly ’best-LLM’ emerged, though CodeT5+ 110M, CuBERT and SPTCode were top-performers. Regression analysis suggests that embedding size, tokenizer vocabulary, and training dataset characteristics are associated with clone detection performance. To address the second question, this paper explores the ensembling of selected LLMs to improve effectiveness. Results suggest the importance of score normalization and favoring ensembling methods like maximum or sum over averaging. Also, findings indicate that ensembling approaches can be statistically significant and effective on larger datasets, where the best-performing ensemble improved performance by 37% over its individual LLMs on the commercial large-scale code.

Read PDF

Similar papers

#small language model Preprint Aug 2026

Vulnerable Code Search: Transferable Attack for Code Language Models

This paper introduces a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability and demonstrates that this attack, even when computed using smaller code embedding models, is highly effective and transferable to larger, closed-source embedding models.

Kaicheng Wang, Liyan Huang, Jesse Thomason et al. · 0 citations
Conference Open access 2026

Large Language Model Vulnerabilities

: Large language models are increasingly being deployed in safety-critical domains, yet remain vulnerable to jailbreak attacks that circumvent safety alignments. This systematic review synthesizes empirical jailbreak research published between 2024 and 2025, using a PRISMA-guided search protocol, followed by BERTopic-b...

Meda Račaitytė, Hélder Bastos, R. Ribeiro et al. · 0 citations
Preprint Sep 2026

Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation

Static analysis remains a cornerstone of software security, yet the effectiveness of tools such as CodeQL is often limited by the substantial manual effort required to develop high-coverage query suites. While large language models (LLMs) have emerged as a potential solution for automated code reasoning, their practica...

I. Irsan, Ratnadira Widyasari, Huihui Huang et al. · 0 citations
Preprint Aug 2026

A Unified Model for Cross-Domain Clone Detection via Model Merging

The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmentation crisis in clone detection. Current deep learning detectors are domain specialists that degrade significantly outside their training distribution, with F1 drops excee...

Palash Ranjan Roy, Banani Roy, Kevin A. Schneider et al. · 1 citation
Preprint Aug 2026

MalTotal: Cost-Effective and Language-Agnostic Malicious Code Poisoning Detection for Millions of Repositories

MalTotal leverages LLM-assisted semantic reasoning to identify sensitive APIs, perform hybrid semantic slicing, and reconstruct malicious behavior contexts while reducing analysis overhead, demonstrating the effectiveness, scalability, and cost-efficiency of MalTotal in mitigating large-scale code poisoning attacks.

Jian Zhao, Shenao Wang, Qingyang Wu et al. · 1 citation · ⚡1
Preprint Aug 2026

Detecting Contaminated Code-Generation Prompt Batches via Influence Functions

Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Existing defenses typically rely on predefined threat models or known vulnerability patterns, limiting their effectiveness against novel attacks. We propose CodeSIFT, a thr...

Francesco Quinzan, Noor Munir, Yi-Shun Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.