Does task decomposition improve automatic NLG evaluation?
This work systematically compares LLMaJ methods with and without decomposition on multiple NLG datasets and finds that, when human labels are available, LLMaJ without using task decomposition can perform comparably to human annotators.