A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability
Abstract
Bacterial vaginosis (BV) concerns involve intimate symptoms, stigma, diagnostic uncertainty, medication use, pregnancy, sexual health, and self-care. Publicly accessible generative artificial intelligence chatbots offer immediate and potentially non-judgmental information, but the quality and safety of specific responses remain uncertain. To compare five publicly accessible generative AI chatbot interfaces using standardized, researcher-developed patient-oriented questions about BV. In this study, patient-oriented denotes consumer-facing wording and does not imply direct patient derivation or validation. Sixty-one standardized English-language questions were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn sessions, yielding 305 responses. Five senior obstetrician-gynecologists independently evaluated safety, accuracy, study-specific guideline-anchored concordance, and empathy using a question-specific reference framework. Independent ratings were locked before adjudication and were used to estimate inter-rater agreement. Consensus scores were used for primary comparisons, and a sensitivity analysis used the median of the five locked ratings. Readability was assessed separately using six formula-based indices. Paired comparisons used Cochran's Q, Friedman, McNemar, and Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Thirty-one responses (10.2%) were classified as unsafe or potentially unsafe. Observed interface-specific rates ranged from 6.6% (95% CI, 2.6%–15.7%) to 14.8% (95% CI, 8.0%–25.7%), while the matched binary comparison did not detect an overall difference (Cochran's Q = 2.596, df = 4, P = 0.627). Under the documented query conditions, differences were detected in accuracy (Kendall's W = 0.686), study-specific guideline-anchored concordance (W = 0.243), empathy (W = 0.322), and all six readability indices (W range, 0.482–0.679; all P < 0.001). The sensitivity analysis produced the same inferential conclusions. Unsafe content occurred in every interface and was concentrated in recurrent-BV, sexual-health, medication, and self-care scenarios. This exploratory study characterizes a recorded sample of specific chatbot responses, not stable or repeatable performance characteristics. Clinically relevant risks were observed in every interface, and the non-significant safety comparison should not be interpreted as equivalence. Because each researcher-developed prompt was submitted once and the interfaces were queried on different dates in a fixed order, the findings do not establish a stable ranking of underlying model capability, clinical effectiveness, or suitability for unsupervised care.