Skip to content

Noisy Test-Time Reinforcement Learning for Code LLMs

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

The Noisy Test-time Reinforcement Learning framework (NTRL-Code) is proposed, which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs.

Abstract

Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at https://github.com/Xikai97/NTRL-Code.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-in...

Fan Chen, Ziheng Jiang, Zi-Yun Wei et al. · 0 citations
#machine learning Preprint Sep 2026

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To m...

Jiacheng Xu, Feng Chen, Xiu-Neng Xu et al. · 0 citations
#natural language process... Preprint Oct 2026

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations su...

Shao-Kun Zhang, Yi-Fan Zhang, Jian Hu et al. · 0 citations
Preprint Aug 2026

CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

CHORUS is presented, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves, and consolidates the resulting specialists into a single 4B model.

He-Jia Zhang, Sheng Lu, Zhongming Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

TTRSD separates update direction from update position, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher, demonstrating cross-dataset generalization while preserving inherent reasoning integri...

Shu-Ning Wang, Zhi-Heng Wu, Xun-Lan Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the o...

Naveen Vakada, Ming-Yuan Li, Shao-Xiong Ji · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.