Skip to content

Author

Yanjun Gao

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

Performance evaluation of a large language model for medication management tasks.

PURPOSE Large language models (LLMs) have proven performance for certain diagnostic tasks; however, limited studies have evaluated their consistency in recommending appropriate medication regimens for a given diagnosis. Medication management is a complex task that requires synthesis of drug formulation and complete order instructions for safe use. Here, the performance of GPT-4o, an LLM available with OpenAI's ChatGPT, was tested on 3 medication management tasks. METHODS GPT-4o performance was tested on 3 medication tasks: identifying available formulations for a given generic drug name, identifying drug-drug interactions (DDIs) for a given medication regimen, and preparing a medication order for a given generic drug name. For each experiment, the model's raw text response was captured exactly as returned and evaluated using clinician evaluation in addition to standard LLM metrics, including Term Frequency-Inverse Document Frequency (TF-IDF) vectors, normalized Levenshtein similarity, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE-1/ROUGE-L) F1 score between each response and its reference string. RESULTS For the first task of drug-formulation matching, GPT-4o had 49% accuracy for generic medications being matched to all available formulations, with an average of 1.23 omissions per medication and 1.14 hallucinations per medication. For the second task of drug-drug interaction identification, the accuracy was 54.7% for identifying the DDI pair. For the third task, GPT-4o generated order sentences containing no medication or abbreviation errors in 65.8% of the cases. CONCLUSION Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.

Kelli Henry, Steven Xu, Kaitlin Blotske et al. · 0 citations
#machine learning Preprint Jan 2026

Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues

It is found that sensitivity to demographic cues is distributed across internal units and varies substantially across models, suggesting that effective mitigation requires understanding distributed, task-specific mechanisms rather than manipulating a small set of identified neurons alone.

Shiyue Hu, Ruizhe Li, Yanjun Gao · 1 citation · ⚡1
#natural language process... Preprint Aug 2026

Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models

A systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs finds that at the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior.

Saksham Khatwani, He Cheng, M. Afshar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.