Large Language Model Semantic-Guided Counterfactual Data Augmentation for Cross-Domain Sequential Recommendation
Abstract
Cross-domain sequential recommendation leverages source-domain interactions to alleviate target-domain sparsity, yet existing methods struggle with semantic gaps, insufficient training signals under extreme sparsity, and negative transfer caused by user heterogeneity. To address these issues, this paper proposes LLM-CFCDR, a cross-domain sequential recommendation framework guided by large language model (LLM) semantics and counterfactual data augmentation. LLM-CFCDR integrates four coupled modules: 1) LLM-guided cross-domain knowledge alignment module that extracts multi-granular item semantics and constructs cross-domain semantic bridges through a graph attention network; 2) counterfactual data augmentation module that performs structured interventions on target-domain sequences and generates preference-consistent virtual samples under LLM semantic constraints and propensity screening; 3) causal-effect-aware sequence encoder that distinguishes factual from counterfactual signals through dual attention and estimates individual treatment effects via inverse propensity weighting; and 4) adaptive transfer mechanism that modulates source-to-target knowledge transfer per user according to causal effect strength and causal entropy, realizing more transfer for strong causality and less for weak causality. Theoretical analysis establishes the asymptotic unbiasedness of the clipped-IPW causal estimator and a Rademacher-complexity generalization bound. Experiments on the Amazon, Douban, and MovieLens-BookCrossing datasets show average improvements of 5.3% and 6.5% over the strongest LLM-enhanced baselines on HR@10 and NDCG@10 (p < 0.01), together with gains of 9.3% and 3.7% on Coverage@10 and Diversity@10 over traditional cross-domain baselines, with the largest benefits observed for cold-start users.