Skip to content

Verifier-Guided Recombination Search for Token-Efficient Test-Time Compute in Countdown

· 0 citations · 6 references

TL;DR

Verifier-Guided Recombination Search is introduced, a post-hoc inference strategy that operates entirely on already-generated Best-of-N rollouts that aggregates correct reasoning fragments from all rollouts, enabling the construction of solutions that did not appear in any individual generation.

View source

Similar papers

Jul 2026

Test-Time Scaling via Error Localization

This work introduces Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization and establishes strictly dominating Pareto frontiers across sequential reasoning domains.

Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta et al. · 0 citations
Jul 2026

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits e...

Sietse Schelpe · 0 citations
Jul 2026

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

This work shows that training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy, and generates answer-blind data, because no correctness filter can see this damage in the data.

Jungseob Lee, Seungyoon Lee, Suhyune Son et al. · 0 citations
#machine learning Review Sep 2026

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is wri...

Fnu Aditi · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.