Grammatical"grandmother neurons"are rare in LLMs
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or"probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds...