Noisy Test-Time Reinforcement Learning for Code LLMs
The Noisy Test-time Reinforcement Learning framework (NTRL-Code) is proposed, which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate pro...