Preprint
Aug 2026
Scaling Inherently Interpretable Language Models
Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.
Guide Labs Team, Andreas Madsen, A. Ismail et al.
· 3 citations
· ⚡1