Jul 2026· ACM Transactions on the Web· 0 citations· 29 references
TL;DR
This work represents the first attempt to develop a fair summary evaluation framework for social media posts and is likely to spark future research in this domain.
Abstract
Text summarization has a wide range of applications, from condensing lengthy documents to summarizing user-generated content (such as tweets, Facebook posts, and Reddit discussions) that reflect diverse viewpoints on socially significant topics. Traditionally, summarization has been considered a reader-centric task, and algorithms are evaluated against human-written reference summaries using metrics such as the ROUGE score. However, this approach has two primary limitations when applied to user-generated content, especially written on sensitive issues. First, through surveys with participants from various countries, we demonstrate that human-written reference summaries can contain implicit biases. These biases introduce significant variations in the evaluation of summarization algorithms, raising concerns about their reliability in identifying summaries that effectively meet reader expectations. Second, we show that algorithmically generated summaries, as ranked by traditional evaluation methods, often fail to align with authors’ preferences. These issues render the existing summary evaluation framework suboptimal for both readers and the authors of the content being summarized. In this work, by reimagining summarization of user-generated content as a two-sided problem, we emphasize the importance of satisfying both readers and authors. Specifically, we take a two-pronged approach to tackle the issues with existing evaluation setup: (i) we introduce the concept of inequality in reader satisfaction to tackle the implicit biases in reference summaries; and (ii) we propose a novel evaluation framework based on satisfying the authors of the content. More specifically, we present an author satisfaction-based evaluation metric CROSSEM which, we show empirically, can complement the current summary evaluation paradigm. This work represents the first attempt to develop a fair summary evaluation framework for social media posts and is likely to spark future research in this domain.
News organizations introduce bias into their coverage via the choices they make about which topics to cover (or ignore) and how to frame the issues they do decide to cover. Here, we introduce the Media Bias Detector, a scalable computational framework that integrates large language models (LLMs) with near-real-time new...
Samar Haider, Amir Tohidi, Jenny S. Wang et al.· Science Advances· 0 citations
A novel abstractive summarisation framework designed to distill coherent and semantically rich summaries from social media discussions that consistently outperforms mainstream summarisation methods, including advanced systems like ChatGPT, particularly in preserving semantic alignment and improving readability under no...
A. Papagiannopoulou, C. Angeli· Expert systems· 0 citations
An expert human evaluation is conducted, measuring summary preferences based on information satisfaction given a specific person’s background and use case, and finds that both traditional and LLM-based metrics are insufficient measures of information satisfaction and agree poorly with human judgment.
Isabel Cachola, W. Walden, Reno Kriz et al.· Proceedings of the 2026 ACM...· 0 citations
This work reimagines recommender systems not merely as engines of engagement, but as accountable infrastructures that uphold democratic values that affect both individuals and society at large.
Aishwarya Satwani· Proceedings of the 20th ACM...· 0 citations
The results demonstrate that human-curated content-based recommendation can positively and significantly impact readers’ knowledge retention and show that a fine-grained coreference system can approach said level of human curation better than state-of-the-art document retrieval methods.
Florian Debaene, Loic De Langhe, Orphée De Clercq et al.· International Conference on...· 0 citations
ParliamentRAG is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions that addresses risks of dominance of the most frequent speakers, inability to weight speakers accord...