Skip to content
Book Open access

CRED: Calibrated Relational Enhanced Distillation for LLM-Based Pointwise Reranking

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 4311-4315 · 0 citations · 30 references
Computer Science

TL;DR

CRED (Calibrated Relational Enhanced Distillation) is proposed, which integrates Adaptive Teacher Calibration (ATC) to calibrate teacher predictions and amplify score margins, while employing Preference Relation Alignment (PRA) to align the distributional patterns of relevance score differences, enabling the student to capture precise ranking structures.

Abstract

In document reranking, rerankers based on Large Language Models (LLMs) demonstrate superior performance but are constrained by high memory consumption and latency. To develop lightweight yet high-performance LLM-based pointwise rerankers through knowledge distillation, we identify two critical limitations: teachers often yield over-smoothed and inaccurate supervision on hard negative samples, thereby hindering the student's optimization; furthermore, traditional methods underutilize the relevance score differences between candidates, which are crucial for ranking tasks. To address these challenges, we propose CRED (Calibrated Relational Enhanced Distillation), which integrates Adaptive Teacher Calibration (ATC) to calibrate teacher predictions and amplify score margins, while employing Preference Relation Alignment (PRA) to align the distributional patterns of relevance score differences, enabling the student to capture precise ranking structures. To support this approach, we also construct FineDistill, a dataset of 1M samples providing fine-grained score supervision. We distill an 8B teacher into a 0.6B pointwise student. Extensive experiments on TREC and BEIR benchmarks show that our model outperforms leading baselines in both performance and generalization.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

CORE is proposed, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking and compares contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budg...

Tingyu Song, Mingxin Li, Yanzhao Zhang et al. · 0 citations
Jul 2026

Distilling large language models for code generation via ranking supervision.

This work proposes a distillation approach based on ranking supervision that consistently outperforms supervised fine-tuning as well as FKL and RKL baselines in Python code generation, multilingual generation, and data-science scenarios and offers guidance for future research in model compression.

Zhe Ding, Hui Ji, Su Pan et al. · 0 citations
Preprint Aug 2026

GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into...

Qi Qin, Jia-Jie Zhu, Dali Chen et al. · 1 citation
#artificial intelligence Preprint Sep 2026

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

Control comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings.

Vignesh Prabhakar, Jiagi Pan, Anil Babu Ankisettipalli · 0 citations
Preprint Aug 2026

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

A practitioner's study of how to make distillation training efficient is presented, organised around two systems contributions, and a fused, chunked KL loss is introduced, making peak memory linear in the sequence length.

Bakbergen Ryskulov, Iker García-Ferrero, David Montero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.