Aug 2026· Proceedings of the 2nd ACM SIGPLAN International Workshop on Language Models and Programming Languages· 0 citations· 59 references
Computer Science
TL;DR
InvPT applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs with invariant-contrast pairs for positives of varying difficulty.
Abstract
Encoder-based code representation models remain widely used for discriminative tasks such as clone detection and code classification because their small size and low inference cost matter. Yet on invariant programs, semantically equivalent code with different syntax, their representations degrade substantially despite unchanged behavior. We measure this gap across four encoder baselines, two tasks, and four datasets, then test a minimal code-only continued pretraining recipe that consistently recovers part of it. Invariant pretraining (InvPT) applies semantics-preserving transformations and combines masked language modeling with multi-positive supervised contrastive learning. All augmentations of a source function are positives: self-contrast pairs use the same code with different masks, whereas invariant-contrast pairs use transformed code, providing varied difficulty without paired natural-language data. Across all model-dataset comparisons, InvPT improves robustness on transformed test sets by a median of 8.1 percentage points for clone detection (up to 11.0) and 3.6 points for code classification (up to 19.2), while matching or improving standard performance. Ablations identify multi-positive invariant contrast as the main source of these gains. Because the test transformations are composed from the pretraining operator family, the results establish invariance only to that family, not robustness in general. Thus, the contribution is not a new objective but a measurement of where encoder robustness breaks and how much of it a simple code-only recipe recovers.
Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-a...
K. Paulsen, Florian Tambon, Mike Papadakis et al.· 0 citations
This work proposes ‘CodeLite’, a relatively lightweight framework that integrates moderately sized code models, such as CodeBERT, with traditional frequency-based techniques, achieving better performance and computational efficiency compared to larger models.
Aman Swaraj, Sandeep Kumar· ACM Transactions on Software...· 0 citations
The preliminary study reveals that models finetuned for function-level binary code similarity exhibit substantially better instruction alignment than their pre-trained model, suggesting a strong correlation between instruction alignment and function-level embedding quality.
LWVIC4Code is proposed, a non-contrastive representation learning approach specifically designed for Type-IV clone detection that achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and...
Luciano Marchezan, Kévin Delcourt, Eugene Syriani et al.· 0 citations
Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patt...
Zhengyang Shan, Jia-Yu Xin, Yan-Jun Lin et al.· 0 citations
DuaLoc combines two pre-trained language models: UniXcoder for the semantic understanding of source code and GraphCodeBERT for awareness of data-flow structure and fine-tuned with a contrastive objective that shapes the embedding space around the localization task.
Amany AlBatlaa, M. Abdullah-Al-Wadud· Electronics· 0 citations