Derivation boundaries: A structural source of faithfulness failure in large language models
Faithfulness to supplied information is not uniform in large language models: some rules, thresholds, or features shape outputs correctly while others do not, with nothing in a model's own explanations distinguishing the two. An industrial classification task motivates this observation, where supplying a classifier's e...