Preprint
Jul 2026
Pretraining Data Can Be Poisoned through Computational Propaganda
This work demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.
Victoria Graf, Hanna Hajishirzi, Noah A. Smith et al.
· 0 citations