Misinformation without liars: How invisible data-processing choices manufacture false conclusions in automated text analysis
Abstract
Official statistics increasingly draws on non-traditional sources, such as social-media posts and other web-scraped text, analysed with AI language models for nowcasting, indicator production, and crisis monitoring. We show that such pipelines can manufacture false statistics with no misinformation actor involved: ordinary, invisible data-processing choices are enough. Using standard, publicly available zero-shot classifiers, we give fully reproducible demonstrations on conflict-related messages. A routine “N comments” interface string captured by scrapers can collapse a model's judgement that a message concerns armed conflict from near-certain to below the retention threshold, while a benign sentence of equal length does not; the effect is large on a widely used model yet negligible on others. Four equally defensible phrasings of a single analytical question yield almost disjoint datasets from identical inputs, and swapping the classifier alone moves an unambiguous conflict report from retained to discarded. These distortions are silent and idiosyncratic: pipelines return confident, plausible numbers, and the direction and size of error depend on the specific model, prompt, and preprocessing. We argue that interpretation artifacts are a distinct threat to data quality and public trust, and propose a short, demonstrated audit that statistical agencies can adopt.