Skip to content
Preprint

COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

COMPAS (Code-generation Optimization over Models, Prompts, And Decoding Settings), a difficulty-aware method that learns group-specific quality-cost fronts through low-cost model selection and joint prompt-decoding search, then routes each test task to its matching front online without further search.

Abstract

Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one configuration for all tasks, routers choose only a model, and prompt optimizers keep the model and decoding settings fixed. This leaves their joint, group-specific interactions unclear. We therefore examine how these choices interact and observe that prompts and decoding settings interact, tuning effects vary by model, and the best configuration varies by task difficulty. Guided by these observations, we introduce COMPAS (Code-generation Optimization over Models, Prompts, And Decoding Settings), a difficulty-aware method that learns group-specific quality-cost fronts through low-cost model selection and joint prompt-decoding search, then routes each test task to its matching front online without further search. Under a matched search budget on LiveCodeBench, COMPAS improves pass@1 from 45.9% for the best baseline to 52.8% while reducing cost from $36.57 to $4.92. This also transfers to repository-level code generation on SWE-bench, resolving 76.0% of tasks versus 70.0% for the best baseline. Code and the reproducibility artifact are available at https://github.com/gjz78910/COMPAS.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and cont...

Tian-Yu Chen, Yasi Zhang, Rui-Yi Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Direct Optimization of Generators for Search in Automated Theorem Proving

This work extends Compute-Aligned Training to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses and introduces a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search.

Adam Ousherovitch, A. Tewari · 0 citations
Preprint Sep 2026

Quality over Quantity: Diversity-Aware Data Selection for Efficient Verilog Code Generation

Large Language Models (LLMs) have shown remarkable potential in Verilog code generation, yet existing datasets contain con siderable noise and redundancy. Prior data selection methods address only isolated quality aspects, neglect the global diversity of the training set, and cannot capture Verilog-specific structural...

Yi-Heng Shen, Wei Zheng, Xiao Wei et al. · 0 citations
#machine learning Preprint Aug 2026

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throu...

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter et al. · 0 citations
Preprint Aug 2026

EvoMem: Memory-Augmented Evolution for Code Optimization

EvoMem is introduced, a persistent memory architecture for LLM-based evolutionary program search that captures and reuses candidate mutation knowledge and provides evidence that persistent memory can reduce some redundant exploration and improve the reuse and adaptation of successful strategies in LLM-driven evolutiona...

Viktor Volkov, Valentin Khrulkov, Andrey V. Galichin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no...

Subhojyoti Mukherjee, Mahmud Tanjim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.