Skip to content

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.

Abstract

General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.

View source

Similar papers

Aug 2026

CodeLite: A Low-Cost Framework for Code Classification Tasks via Enhanced Code Representations and Lexical Fusion

This work proposes ‘CodeLite’, a relatively lightweight framework that integrates moderately sized code models, such as CodeBERT, with traditional frequency-based techniques, achieving better performance and computational efficiency compared to larger models.

Aman Swaraj, Sandeep Kumar · 0 citations
#artificial intelligence Preprint Sep 2026

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-...

Jin-Hao Zhang, Ze-Yu Liu, Zi-Cheng Yan et al. · 0 citations
#machine learning Preprint Aug 2026

Can LLMs Use Relational Transformer Embeddings?

It is argued that soft-token fusion requires stronger alignment objectives and schema-aware design before it can serve as a reliable route to relational prediction.

Francisco Galuppo Azevedo, Clarissa Lima Loures · 0 citations
#small language model Book Open access Aug 2026

Invariant Pretraining for Robust Code Representations

InvPT applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs with invariant-contrast pairs for positives of varying difficulty.

Yifeng He, Yun-Di Xu, Christopher Castro Gaw Gonzalo et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.