Benevolent Bias in Multi-Turn Human-Agent Dialogue
It is suggested that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate, as off-the-shelf detectors reliably flag overt bias yet largely miss benevolent bias, LLM judges catch more under more explicit detection criteria but increasingly mis classify neutral support as benevolent bias, and demographic context amplifies the false alarms.