Aug 2026· Journal of Studies on Alcohol and Drugs· 0 citations
Medicine
TL;DR
LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review, and benefit from best practices and technical guidelines for LLM utilization.
Abstract
Objective
Policy surveillance typically involves detailed, time-consuming manual screening of policies for inclusion in a final dataset. This screening process involves risks of human error and inconsistent application of inclusion/exclusion criteria, especially in complicated legal landscapes like the US opioid treatment landscape. Large language models (LLMs) could assist human subject matter experts (SMEs) during screening, but LLMs have been understudied for policy surveillance. Therefore, we conducted a test comparing opioid treatment policy screening decisions between SMEs and an LLM.
Methods
Using a Boolean search string in legal software, we identified 99 potentially relevant Massachusetts policies for emergency department opioid addiction treatment, and we downloaded text from government websites. Next, we compared two approaches to screening those policies using pre-defined inclusion and exclusion criteria: (a) manual screening by three SMEs, and (b) an LLM approach. We assessed the overall percentage of inclusion/exclusion decisions where the LLM made the same decision as the SMEs. We also identified the percentage of policies selected for inclusion by the SMEs with which the LLM agreed and potential reasons for discrepancies.
Results
The LLM made the same decision for 96 of 99 policies (97% of the time). All policies that SMEs chose to include (n=2) were also included by the LLM. Discrepancies reflected implicit inclusion and exclusion criteria used by SMEs but not provided to LLMs.
Conclusion
LLMs could serve as a quality control check during opioid policy surveillance research, supplementing human review. The policy surveillance field would benefit from best practices and technical guidelines for LLM utilization.
Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.
D. Dobin, A. Witmer, F. Sweeney et al.· medRxiv· 0 citations
Abstract
Medication errors remain among the most significant preventable causes of patient harm worldwide, contributing to increased morbidity, mortality, prolonged hospitalization, and escalating healthcare expenditure. Large Language Models (LLMs) — including GPT-4, Gemini, Claude, and Llama — have emerged as promising clinical decision-support tools capable of interpreting complex medical terminology, analyzing prescriptions in real time, and flagging potential errors before medications reach the patient. This review synthesizes current evidence on the application of LLMs in prescription error detection, presents a consolidated system architecture and operational workflow for LLM-enabled medication safety pipelines, and critically examines their benefits, limitations, and future trajectory. Evidence from recent clinical evaluations indicates that LLM-based decision-support tools can achieve high concordance with expert pharmacist judgment and measurably reduce near-miss medication events when deployed with appropriate safeguards. However, challenges including AI hallucination, data privacy, algorithmic bias, regulatory ambiguity, and the continued necessity of human oversight must be addressed before widespread clinical adoption. The review concludes that LLMs hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
Keywords: Large Language Models; Prescription Error Detection; Medication Safety; Clinical Decision Support; Artificial Intelligence in Healthcare; Electronic Health Records; Pharmacovigilance
K. T. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.
Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong· Diagnostics· 0 citations
BACKGROUND
We evaluated the reliability of large language models (LLMs) for abstract screening under real-world review practices and qualitatively characterized model-human discordance to inform safe workflow integration.
METHODS
We evaluated GPT-4.0, GPT-5.0, and GPT-5.0-mini on two curated systematic review datasets representing contrasting topic densities, defined by the target-to-background ratio (TBR): a core-subject dataset (TBR 69%) in which the target intervention was central, and a peripheral-subject dataset (TBR 4.4%) in which the target intervention was incidental. Final inclusion after full-text review served as the reference standard. We developed a qualitative taxonomy of disagreements, classifying false negatives as intended human leniency, gray-zone ambiguity, or true LLM misses, and false positives as implicit or additional human exclusion rules, gray-zone ambiguity, or nominal inclusions that increase workload only.
RESULTS
GPT-5.0-mini achieved the best sensitivity-efficiency trade-off (core-subject: 91% sensitivity with 96.7% workload reduction; peripheral-subject: 83% sensitivity with 92.7% workload reduction) and negative predictive value >99% in both datasets. Disagreement was lower when relevance was central (core-subject: 1.6%, 7/430) with no true LLM misses (0/430). In the peripheral-subject dataset, disagreement was higher (10.6%, 74/696), driven mainly by intended human leniency among false negatives (52/56) and gray-zone ambiguity among false positives (12/18), while true LLM misses remained rare (0.4%, 3/696).
CONCLUSION
Many model-human disagreements reflect topic- and workflow-dependent screening conventions rather than intrinsic model failure. LLM-assisted screening may improve efficiency without compromising reliability when accompanied by appropriate safeguards for ambiguous records.
K. Lee, Hakyoung Kim, Dae Sik Yang et al.· Journal of Epidemiology· 0 citations
Introduction Large language models (LLMs) are entering drug-safety and regulatory workflows, yet their behavior at the boundary between unavailable evidence and risk reassurance remains poorly characterized. Drug-induced liver injury (DILI) is a stringent setting because evidence is fragmented across labels, case reports, mechanistic studies, and curated knowledge bases, while unsupported low-risk reassurance can be consequential. Methods We evaluated five LLMs on a fixed 47-drug DILI risk-assessment panel using 1,410 parsed responses from closed-book answering and an evidence-gated protocol that restricted responses to supplied PubMed-derived evidence. Curated DILI resources were excluded from prompts and used only for evaluation. After filtering, 38 drugs had no direct DILI decision-support evidence in the supplied packet, and 9 had direct DILI-relevant evidence; a post hoc PubMed title/abstract recall stress test identified five additional drugs with recoverable direct DILI evidence outside the packet. Results In the no-direct-evidence slice defined by the supplied packet, closed-book models rarely abstained, with drug-level abstention ranging from 5.3% to 33.3%; the evidence-gated protocol required abstention, which all models followed for every no-direct-evidence drug. The same pattern held for recent or low-recognition drugs, where evidence-gated abstention reached 92.0% to 100.0% vs. 8.0% to 49.3% under closed-book answering. Closed-book models also produced high-confidence low-risk responses for DILI-positive drugs, a label-discordant pattern largely removed by evidence gating. Independent expert review of selected responses showed that label discordance did not always imply a clinically unreasonable low-risk category, but identified unsafe reassurance through overconfident wording and under-cautious responses in selected cases. When direct DILI evidence was provided, all models preserved citation-grounded non-abstaining answers. However, they differed in how often they committed to a conclusive rather than an uncertain risk category. Citation-bearing evidence-gated responses cited only the supplied PubMed identifiers and achieved 91.2% to 100.0% concordance with the supplied grade. Discussion These findings identify unsupported reassurance as measurable evidence-boundary behavior in LLM drug-risk assessment and establish a reproducible framework for auditing adherence to an externally supplied evidence boundary, defined by PubMed evidence classification and enforced by prompt policy rather than inferred independently by the model.
Jinwan Shi, Yong Ma, Yinhui Liu et al.· Frontiers in Public Health· 0 citations
: Large language models can generate fluent and often convincing answers in medical contexts, but in high-risk settings, fluency alone is not enough. This paper presents a policy-guided LLM pipeline for medication-related clinical decision support, designed to classify prompts into safety-oriented decision categories: ACCEPT, WARN, DEFER, ESCALATE, or REFUSE. The system combines rule-based risk detection, guardrail routing, and a decision policy layer to handle prompts involving dosage, pregnancy, drug interactions, self-adjustment, and other safety-critical contexts. The pipeline was evaluated on a benchmark of 55 medication-related questions using two baseline models. Both models produced the same overall system accuracy of 60.0%, with a false accept rate of 16.4%, suggesting that the main limitations are not model-specific but structural. Three recurring failure modes emerged: false acceptance of implicitly risky prompts, over-escalation of educational or professional-context questions, and under-escalation of self-adjustment or dangerous-intent cases. These findings make the system useful not only as a prototype, but also as a transparent framework for studying where safety-oriented LLM pipelines succeed and where they still fail.
Ana Stevanović, Mlađan Jovanović· SINTEZA· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 17, 2026
A USAF cadet and a Lincoln Laboratory researcher found AI chatbots can help nontechnical service members produce viable software applications for their unique problems.