Jun 2026
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
This work introduces a persona-driven rewriting pipeline that conditions user turns on low agreeableness and pairs this with warm, de-escalating assistant responses, and shows that safer empathetic fine-tuning is achievable through data design alone, without safety labels, harm detectors, or changes to the training objective.
A. Cheung, Yi Yang
· arXiv.org · 0 citations