Fair Like Us? Auditing LLM Alignment in Resource Allocation
It is found that large language models tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.
Qi-Shen Han, Hadi Hosseini, Joshua Kavner et al.
· 0 citations