Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test
Controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in the authors' runs.