#natural language process...
Aug 2025
Evaluating Style-Personalized Text Generation: Challenges and Directions
This work critically examines the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges, and finds strong evidence that employing ensembles of diverse evaluation metrics consistently outperforms single-evaluator methods.
Anubhav Jangra, Bahareh Sarrafzadeh, Adrian de Wynter et al.
· arXiv.org · 2 citations