Sep 2026· Proceedings of the 20th ACM Conference on Recommender Systems· 0 citations· 5 references
TL;DR
It is found that LLM judges are influenced by descriptions of a recommendation algorithm’s optimization objective, even when evaluating identical recommendation outputs, a phenomenon the authors term intent-description anchoring bias.
Abstract
Large Language Models (LLMs) are increasingly used to evaluate recommendation systems, but are known to exhibit systematic biases. We study whether LLM judges are influenced by descriptions of a recommendation algorithm’s optimization objective, even when evaluating identical recommendation outputs. We find that they are, a phenomenon we term intent-description anchoring bias, where an algorithm’s stated objective influences judgments beyond what is supported by the recommendations themselves. Across four frontier LLMs from three commercial providers, providing algorithm descriptions in the evaluation prompt increased diversity scores for identical recommendations by up to 0.89 points (Cohen’s d = 1.82, p < 10− 18), with substantial variation across models. Our results show that contextual information unrelated to recommendation quality can bias LLM-based evaluation, motivating protocols that hide algorithm metadata from LLM judges or apply mitigation strategies.
The utility of LLMs in selecting an effective explanation method for a given application is studied and four practical recommendations are derived: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy.
Kathrin Wardatzky, Oana Inel, Luca Rossetto et al.· Proceedings of the 20th ACM...· 0 citations
Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice...
It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.
A. Kapetanović, Kemal Altwlkany, Andro Merćep et al.· 0 citations
This work reimagines recommender systems not merely as engines of engagement, but as accountable infrastructures that uphold democratic values that affect both individuals and society at large.
Aishwarya Satwani· Proceedings of the 20th ACM...· 0 citations
When evaluating top-N recommendations, objectives beyond accuracy are increasingly considered, reflecting a broader shift toward user-centric evaluation. However, existing evaluation approaches present several limitations: (1) they fail to quantify how much users actually value beyond-accuracy objectives, (2) they lack...
T. Zada· Proceedings of the 20th ACM...· 0 citations
Large language models (LLMs) are increasingly used for product recommendation, but evaluating their recommendations presents challenges that differ from conventional information retrieval and recommender systems. LLMs can generate recommendations without an explicit candidate set, and repeated responses to the same que...
E. Malthouse, Kun-Yu Lee, Jing Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.