Aug 2026· Proceedings of the 2026 ACM Conference on Human-AI Complementarity and Alignment· 0 citations· 54 references
Computer Science
TL;DR
This work proposes a two-stage framework to assess the severity of false claims during disasters, and investigates false claim severity assessment as a human-AI alignment problem, evaluating whether models can reproduce human judgments under a shared evaluation rubric rather than merely predicting severity labels.
Abstract
False information spreads rapidly on social media during disasters and can undermine emergency response efforts, public trust, and crisis communication. Existing research primarily focuses on determining whether social media posts contain false information, but provides limited insight into the specific false claims embedded within posts and the severity of individual false claims. To address the limitations, we propose a two-stage framework to assess the severity of false claims during disasters. In the first stage, we develop a false claim extraction agent that identifies false claims from multimodal social media posts containing text, images, videos, and links. A subsequent verification step validates extracted claims with supporting evidence. In the second stage, we define false claim severity as the combination of two complementary dimensions: believability, which determines the likelihood that a claim will be believed, and harmfulness, which captures the potential consequences if it is believed. Human annotators assess both dimensions to construct a claim-level severity benchmark using false claims extracted from Reddit posts related to hurricanes and wildfires. Building upon this benchmark, we investigate false claim severity assessment as a human-AI alignment problem, evaluating whether models can reproduce human judgments under a shared evaluation rubric rather than merely predicting severity labels. Experiments on the benchmark show that traditional supervised models exhibit limited alignment with human judgments, whereas Large Language Models (LLMs) achieve substantially stronger performance. Among the evaluated strategies, in-context learning consistently achieves the strongest alignment with human judgments, highlighting the importance of human examples and shared decision criteria for severity assessment.
Rumors, misleading claims, and other factually risky information units may gain visibility during crises before verification processes are complete. Existing detection systems often address factual status or diffusion separately, whereas operational monitoring requires identifying units that are both factually risky an...
J. E. Solanes, Juan José Climent-Ferrer, Flavio Moriniello et al.· Electronics· 0 citations
Progress now depends less on architectural novelty than on label construction, temporally honest evaluation and outcome measurement beyond the confusion matrix, and almost no study measures downstream effects on audiences or on fact-checking workflows.
Mary Oluwakemi Abioye· Asian journal of mathematics...· 0 citations
Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial intelligence (AI) can produce plausible messages that resemble eyewitness reports. This study investigates whether text-based AI detectors can reliably distinguish human-a...
False and misleading information on social media during broadscale societal crises, or so-called mega-crises, is a growing concern across academic fields. While previous studies have offered valuable insights, research remains fragmented.
This interdisciplinary scoping review summarizes current knowledge by...
Sofia Johansson, P. Rodin, Alice Srugies et al.· Frontiers in Communication· 0 citations
Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support conflict verification tasks, we present ContraNote, a large-scale real-world dataset curated from X's Community Notes s...
Shu-Ning Zhang, D. Shi, Bo-Hao Chu et al.· 0 citations
Event Causality Identification (ECI) is a crucial task in knowledge discovery that extracts structured causal relationships between annotated event mentions from unstructured text. However, existing approaches typically rely on extensive labeled data, which is scarce for specialized domains and topics. Although Large L...
Ze-Fan Zeng, Yue-Hang Si, Xing-Chen Hu et al.· Proceedings of the Thirty-Fi...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.