recirculation
We describe an inference-time architectural enhancement for off-the-shelf foundation models that systematically reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency during generation, though it requires serial processing in the prefill phase...