Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing
Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two c...