Skip to content

CodeLite: A Low-Cost Framework for Code Classification Tasks via Enhanced Code Representations and Lexical Fusion

Aug 2026 · ACM Transactions on Software Engineering and Methodology · 0 citations · 60 references

TL;DR

This work proposes ‘CodeLite’, a relatively lightweight framework that integrates moderately sized code models, such as CodeBERT, with traditional frequency-based techniques, achieving better performance and computational efficiency compared to larger models.

Abstract

The prevailing approach in software engineering for code classification tasks is to rely on large-scale code language models such as StarCoder and CodeLlama. While these models have demonstrated strong performance, they come with the price of intensive computations that raise energy concerns and limit their deployment in resource-constrained environments. In this work, we propose ‘CodeLite’, a relatively lightweight framework that integrates moderately sized code models, such as CodeBERT, with traditional frequency-based techniques, achieving better performance and computational efficiency compared to larger models. Our approach comprises multiple stages, starting with the enhancement of encoder-based code models, where we leverage their intermediate-layer information to capture richer lexical and syntactic signals beyond the default [CLS] token. In parallel, we employ a regex-enhanced TF-IDF component tailored for source code to capture complementary frequency-based patterns. These representations are then combined with the learned code embeddings through suitable fusion techniques for final decision-making. Evaluated across multiple downstream tasks, including programming language identification, authorship attribution, and AI plagiarism, CodeLite consistently outperforms strong baseline models, achieving up to 4.2% absolute accuracy improvement over other approaches and 11.8% incremental gain over ablation components while incurring significantly lower training costs. We further conduct several statistical tests and ablation studies to validate the contribution of each component in the proposed framework. Overall, our results demonstrate that a careful combination of lightweight models along with task-aware feature enhancements can serve as a practical and efficient alternative to heavyweight language models for various software engineering tasks.

View source

Similar papers

#small language model Book Open access Aug 2026

Invariant Pretraining for Robust Code Representations

InvPT applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs with invariant-contrast pairs for positives of varying difficulty.

Yifeng He, Yun-Di Xu, Christopher Castro Gaw Gonzalo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Retrieval-Augmented Generation for Scientific Code Understanding

The results indicate that front-loading code understanding into a reusable, codebase-specialised vector store enables small local models to deliver grounded and repository-specific answers, making the agent well suited as a privacy-preserving development tool for in-house scientific codebases.

Aaron Nobile, Andreas Adelmann, Mohsen Sadr · 0 citations
Preprint Aug 2026

CodeHID: Learning an Addressable Hierarchical Code Index for Generative Code Retrieval

This paper proposes CodeHID, a generative code retrieval framework that reformulates the code retrieval task from flat candidate matching to coarse-to-fine semantic address generation and outperforms existing sparse retrieval, pre-trained code models, dense code retrieval, and generative retrieval baselines by a large...

Zhen Li, Yuhong Chen, Wenhao Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-a...

K. Paulsen, Florian Tambon, Mike Papadakis et al. · 0 citations
Conference Sep 2026

Comparative Analysis of Rewriting-Based Large Language Model-Generated Synthetic Code Detection

The case of synthetic code detection (i.e., written by a human or generated by AI) becomes more important in the field of education and security. Current code detection tools that depend on probabilistic and classification-based approaches have limitations in detecting various programming styles. To address this challe...

Maulana Arya Alambana, Arif Nurwidyantoro · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.