Skip to content
Book Open access

Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM Reasoning

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 5951-5962 · 0 citations · 49 references

TL;DR

This work proposes Step-wise Training for In-context Reasoning (STIR), a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query.

Abstract

Large language models (LLMs) have shown strong performance in mathematical reasoning, supported by approaches such as In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG). However, existing methods often provide problem-level examples, which is too coarse-grained for multi-step reasoning to cause informational redundancy, and structural misalignment. To address this limitation, we propose Step-wise Training for In-context Reasoning (STIR) to provide step-synchronized and logically targeted guidance to enhance the model's mathematical reasoning capabilities. STIR enables a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query. First, We decompose expert solutions into Step-Level Reasoning Units inspired by human thinking patterns. Leveraging this data, a Step Retriever is trained for logical continuity to map current reasoning states to relevant subsequent steps. Then a Step Reasoner is trained to decide when to retrieve tailored step examples and incorporates this guidance into reasoning. We further extend STIR with a Process-aware Reinforcement Learning phase using Group Relative Policy Optimization to learn to self-formulate search queries and optimizes the decision-making policy. Experiments on seven benchmarks demonstrate that STIR achieves accuracy improvements ranging from 1.86% to 17.96%, maintaining lower token efficiency than baselines. Analysis via our proposed DSM, TCN and RCR metrics shows that STIR improves reasoning capability, achieving DSM scores ranging from 4.54 to 29.87 across backbones and significant improvements over the baseline in both TCN and RCR.

Read PDF

Similar papers

Preprint Aug 2026

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

DynaRule is proposed, an end-to-end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step-wise process, and can re-attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi-step reasoning.

Bo-Han Yu, Pengfei Cao, Chen Han et al. · 1 citation
Preprint Sep 2026

Knowledge-as-Skill: A Structural Design for Autonomous Knowledge-Base Use by LLM Agents

Retrieval-augmented generation (RAG) gives large language models (LLMs) access to external knowledge, but its conventional retrieve-concatenate-generate pipeline makes retrieval decisions on behalf of the model. As tool use and agent loops become more reliable, an agent can decide whether to retrieve, what to inspect,...

Jiang-Xu Wu · 0 citations
#natural language process... Preprint Aug 2026

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

INSPIRE is an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-orie...

Shuai Wang, Jiayi Kuang, Ying-Hui Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability

Recently, Large Language Models (LLMs) have gained significant attention due to their strong language understanding and generation capabilities, demonstrating impressive reasoning abilities as well as effective utilization of external knowledge. Many studies have proposed methods that specialize in improving performanc...

Keuntae Kim, Eunhye Jeong, Yongsuk Choi · 0 citations
Preprint Aug 2026

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

SkillReason-Bench is introduced, a large-scale cross-domain benchmark containing 3,729 queries and a retrieval corpus of 61,228 skills spanning nine domains and SkillRea- son is proposed, a two-stage framework that uses chain-of-thought rea- soning as training-time supervision for skill retrieval.

Donghong Jiang, Endian Lin, Luoping Cui et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.