Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions.
False expertise claims were the most vulnerability-inducing pressure pattern, whereas emotional escalation was associated with the highest scores, and adversarial caregiver pressure was associated with higher than lower Safety Composite Score values for all four models across all ten topic categories and severity level...