ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
This work proposes ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing, and introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation.