Parsing Causal Relationships in Social Science Publications
Abstract
Understanding the causes of phenomena is a key goal of scientific enquiry. For historic reasons, social scientists are reluctant to explicitly discuss causal assumptions. Nevertheless, causal claims do appear in the literature. Systematically cataloging these causal claims and parsing them into structured cause-effect relationship would allow us to synthesize them and, ultimately, construct theories based on oft-repeated causal claims. However, manually extracting these claims is infeasible due to the scale of the literature and linguistic ambiguity, and existing automated methods often lack domain specificity or fail to extract specific cause-effect pairs. To overcome these limitations, we introduce a BERT-based multitask deep learning model trained on a novel, domain-specific benchmark dataset of 3,014 annotated sentences from social science papers. The model simultaneously (1) identifies sentences containing causal claims, (2) extracts cause and effect spans, and (3) links the associated causal pairs. In our study, the multi-head model outperformed a model with sequential architecture, as well as two generative LLMs (Llama 3 8B and Qwen3 8B), which were prompted to code the sentences with a few examples of the coding schema. Specifically, the multi-head model had the highest macro F1-score across all three subtasks and lower false-positive error propagation compared to the sequential model. Its agreement with human coders on the held-out test set (Krippendorff’s α = 0.81) was of similar magnitude to the interrater agreement among the human coders, computed on an interrater reliability subsample (α = 0.80).