Skip to content

Author

Joji Suzuki

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Evaluating Large Language Models in Response to Questions on Substance Use: Helpful or Harmful?

BACKGROUND Individuals with substance use disorders (SUD) are obtaining health-related information from various large language models (LLMs). We aimed to assess whether LLMs provide responses concordant with the current evidence base and whether they provide harmful responses. METHODS Twenty questions related to SUD were posed to three LLMs (Gemini-1.5-pro-001, Claude-3-5-sonnet, and GPT-4) in May 2024. Each response was independently rated by three experienced addiction specialists, and disagreements were resolved by two additional experienced addiction specialists. All raters were blinded to the LLM. Each rater assessed whether (I) a competent addiction specialist would agree with the response, (II) the response contained stigmatizing language as defined by National Institute on Drug Abuse, or (III) the response contained harmful content. RESULTS 88% of responses were rated as competent and 92% as not harmful. Gemini-1.5-pro-001 had the highest rate of competence (95%), followed by Claude-3-5-sonnet and GPT-4 (both 85%). Gemini-1.5-pro-001 produced no harmful responses, while Claude-3-5-sonnet and GPT-4 produced 10% and 15%, respectively. 30% of responses from both Gemini-1.5-pro-001 and Claude-3-5-sonnet had contained stigmatizing language, compared to 10% for GPT-4. CONCLUSIONS While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.

Samuel Maddams, Shan Chen, Danielle S. Bitterman et al. · 0 citations