Skip to content
Open access

Evaluation of Large Language Models for an AI Chat Assistant Focused on Pumas and Pharmacometrics

2026 · Quantitative Medicine · Vol 1 · 0 citations · 18 references

TL;DR

Overall, it is concluded that model selection for domain RAG applications should be treated as a modular process that considers trade-offs between metric weights and insights from embedding-based clustering, so AskPumas can adapt its priorities as the LLM landscape evolves.

Abstract

The Large Language Model (LLM) landscape is constantly changing, with new models emerging rapidly and outperforming established benchmarks. For professionals working on LLM applications, this means constantly being aware of top-performing LLMs that can replace current LLMs in their internal AI applications; an LLM that is considered the optimal choice in a scientist’s current AI architecture cannot be assumed to remain the optimal choice as models constantly evolve. At PumasAI, while building AskPumas, a Retrieval Augmented Generation (RAG) based AI chat assistant focused on Pumas and pharmacometric-related inquiries, we noticed the importance of finding a structure to assess numerous LLMs. The motivation behind this research is to establish a framework to evaluate LLMs and select and rank them based on different criteria. This paper will explore the methodologies of our evaluation criteria, how this has served in helping select the ideal LLM for AskPumas, and how our established framework can serve as a guide for developers who seek to find the optimal LLM for their domain-specific AI application. Using this framework, we identified gpt-5-chat , gemini models, qwen3-max , and grok-4-fast as top-performing LLMs for AskPumas. The differences among the highest-ranked models are small (~1-2%), suggesting that we identified sets of strong candidates for AskPumas, rather than a single definitive winner. Overall, we conclude that model selection for domain RAG applications should be treated as a modular process that considers trade-offs between metric weights and insights from embedding-based clustering, so AskPumas can adapt its priorities as the LLM landscape evolves.

Read PDF

Similar papers

Have Large Language Models Improved Research Methodology?

Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.

Fredric Narcross, Robert Marks · 0 citations
Preprint Jul 2026

Knowledge Distillation for Automated AI Tutor Evaluation

The rapid integration of Large Language Models (LLMs) into K-12 and higher education has outpaced the development of reliable methods for evaluating their pedagogical quality. As the research community starts to explore the space of automating evaluation of AI tutors, we introduce FATE (FLC AI Tutor Evaluator), a specialized 8B-parameter language model designed to evaluate AI tutors. Aligned with the four core evaluation tracks from the BEA 2025 Shared Task, our model assesses pedagogical ability across Mistake Identification, Mistake Location, Guidance, and Actionability. Because pedagogical evaluation is a specialized task with limited labeled data, we leverage knowledge distillation from a frontier LLM to generate additional supervision, yielding absolute performance gains up to 22.63 percentage points. Finally, we demonstrate FATE's utility as an automated evaluator by benchmarking instructional responses generated by popular commercial models, including ChatGPT, Claude, Gemini, and DeepSeek. On average, we have found that Gemini 2.5 Flash perfomed best (82.88%), then ChatGPT 5.5 Instant (80.75%), followed by DeepSeek V4 Flash (80.13%) and Claude Sonnet 4.6 (74.00%).

Tahmid Al Hannan, Diego García, Alex K Njoroge et al. · 0 citations
Open access Aug 2026

Human VS AI: comparison of scientific paper drafting capabilities

Large language models (LLMs), a type of artificial intelligence (AI), are increasingly popular tools used for everyday and work-related activities and tasks. Their application in medicine is widely researched and has been used recently to help write scientific papers in various fields. LLMs can draft sections of manuscripts or whole papers far more quickly than human writers. However, they need appropriate prompting to draft near-complete and worthwhile papers. In the current study, we used two different LLM models: ChatGPT-4o and Claude 3.5 Sonnet, to test AI’s scientific writing capabilities. We took an article written by the authors of the current study on the topic of fluorescent cholangiograms and ran it through the models to help build prompts for the article to be written by the AI. After that, we loaded the data and references used for the article into both AI tools and prompted them to write sections of the article (or a whole article if possible) on the same topic while allowing for independent choice of statistical analysis. The results of the human-written article and the AI-generated ones were compared, evaluating the information used from the references, the types of statistical analysis methods used, the conclusions drawn, and the time it took to complete the task.

T. Yotsov, M. Karamanliev, M. Petkov · 0 citations
Review Open access 2026

LLMs and Generative AI for Everything?

The research field of Natural Language Processing (NLP) has experienced a major shift since the introduction of Large Language Models (LLMs). All facets and application scenarios within NLP have been impacted by the use of LLMs. Current research as well as practice of text processing tools is focused mainly on the application and development of LLMs. Major investments, not only by LLM providers but also other companies applying LLMs in their workflows, have only solidified the role of LLMs in NLP - and in other research and application areas - as part of the artificial intelligence boom in recent years. However, limitations and downsides of the application of LLMs have also emerged. Problems regarding the generated texts as well as the environmental impact of the large-scale use of LLMs are just two of many factors that should be critically analyzed, despite the hype and the prevalence of LLMs for NLP tasks. These restrictions provide the main motivation for this thesis. Traditional models as alternatives to LLMs will be discussed from different perspectives. The characterization of traditional models will be progressively developed as features of alternatives to LLMs will emerge during the course of this thesis. This process will be grounded in experiments, observations and evaluations. Several NLP applications will be presented by surveying the state of the art with neural network-based models such as LLMs as well as the current usage of traditional models. The concrete NLP applications comprise information and relation extraction, text classification, text segmentation, text simplification and text summarization. The first half of this thesis will present the emergence of LLMs contextualized along previous developments within NLP. Characteristics of the selected NLP applications will be collected before a structured literature review will display the prevalence of LLMs regarding each application and will discuss if traditional models are still actively researched. Lessons from domains with long-standing development procedures and processes will also be taken into account to provide a purposeful and structured manner of approaching NLP tasks. A collection of challenges within current NLP will conclude the first half of the thesis, which will serve as motivation for the analysis of experiments and applications of the latter half. The second half of this thesis will present observations and evaluations from use cases, aligned towards the challenges recognized in the first half. Through the analysis of these use cases, benefits of applying traditional models will be collected and supported, in particular through the analysis of a text segmentation use case that is purposefully applied with the lessons drawn from the first half of the thesis in mind. The interpretation of these results will conclude in a discussion on the applicability of traditional models in contrast to LLMs and also give recommendations of both model types for different use cases. Concrete use cases for information extraction, entity matching, text classification and text segmentation will be presented, in which traditional models match or surpass the performance of modern methods. Through improved efficiency as well as enhanced explainability and reproducibility in comparison with neural network-based techniques, these showcases demonstrate the continued relevancy of traditional techniques in today's NLP landscape. Overall, this thesis discusses the role of traditional models in current NLP research and practice, especially in contrast and comparison to modern neural network-based approaches including LLMs. The applicability of modern and less modern techniques is analyzed through a case-based analysis of NLP tasks in a structured and purposeful manner.

Robin Jegan · 0 citations
Open access Aug 2026

Provisioning An Adaptive Model to Analyze Uncertainty and Large Language Patterns for Enhanced Document Re-Ranking

A fundamental change in information retrieving (IR) has been brought about by the quick development of large language patterns (LLMs), which go beyond standard keyword inquiries and ranked outcome lists. Retrieving-Augmented Generation that followed, a more interactive and lively regaining process that incorporates different facets of Accessibility to data into the conversation amongst an individual and the internet engines for searching and exploring, is one of the new interaction forms introduced by LLMs, which are now crucial to the development of IR technologies. We examine the complex effects of LLMs on IR, focusing on three different layers from which they have become essential to the retrieving process: the interaction layer, the structure for obtaining information and a computation pipeline functionality that can leverage a richer meaning representation through sophisticated language patterns, as well as the larger IR ecology. This work introduces a trust based adaptive reranking model- ATM (Adaptive Trust Model)that allocates computational resources according to file level uncertainty. Instead of assigning a fixed number of reranker calls per query, ATM focuses computation only where ranking confidence is low. This concentrate on prejudice, fairness, and ethical considerations in addition to evaluation challenges for the latter. The model gives 15–30% reduction in floating point operations (FLOPs) and up to 20% lower latency while maintaining or improving retrieval precision. To illustrate the influence on one area of study, we point to a few current examples of LLMs being employed in the medical field

Jenny Kalaiarasi.S · 0 citations
Open access Aug 2026

EduAssist: Evaluating a Locally Deployed Large Language Model for Educational Document Summarization

The increasing use of large language models (LLMs) for educational text processing has created opportunities for automatic summarization of lengthy learning materials. However, many LLM-based applications rely on cloud-hosted services, while the performance and computational behavior of locally deployed language models for educational document summarization remain comparatively underexplored. This study presents EduAssist, an experimental framework for evaluating a locally deployed Qwen3:1.7B model for educational document summarization. The model was executed through the Ollama runtime using a chunk-based summarization and consolidation pipeline. Experiments were conducted on 12 English-language educational samples covering topics in data mining, machine learning, artificial intelligence, and knowledge-based systems. Evaluation combined text-reduction and efficiency measures with lexical, semantic, and model-assisted assessment using compression ratio, estimated reading-time reduction, processing time, source-reference ROUGE, BERTScore, and LLM-as-a-Judge. The generated summaries achieved a mean compression ratio of 50.83%, reducing the mean source length from 1,074.17 to 478.42 words and yielding an estimated mean readingtime saving of 2.98 min. Mean ROUGE-1, ROUGE-2, and ROUGE-L scores were 0.4917, 0.2371, and 0.2892, respectively, while the mean BERTScore F1 was 0.8264 and the mean LLM-as-a-Judge score was 7.96/10. Summary generation required an average of 194.33 s per sample. Exploratory analysis further showed that greater compression was associated with lower source-reference lexical retention, whereas source-summary BERTScore F1 values remained comparatively stable across the evaluated samples. Overall, the findings provide exploratory empirical evidence for the feasibility of local educational document summarization using the evaluated Qwen3:1.7B configuration, while highlighting the need for larger datasets, model comparisons, independent reference summaries, and human evaluation.

Mohammed Afzal, Shifa Tahreem · 0 citations