Book
Open access
Jul 2026
When AI Coding Assistants Leak Training Data: A Study of LLM Memorization in Code Generation
It is confirmed that memorization persists in modern LLMs and is influenced more by a complex interplay of training domain, dataset composition, architectural choices, and content characteristics, rather than parameter count alone.
Xiaoyu Cheng, Kundi Yao, Pengyu Nie et al.
· AIware · 0 citations