Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.
Generative artificial intelligence (GenAI) is a large language model (LLM) that has the ability to generate media based on user-provided prompts. Given the demonstrated capabilities of models such as ChatGPT in information synthesis and programming, there is growing interest in their potential role within the research process. However, little work has evaluated recent GenAI models for research tasks in the domain of statistical research. This case study examines GenAI as a tool for developing a literature review and translating methodology from academic papers into code, for the topic of dynamic treatment regime (DTR) estimation via the dynamic weighted ordinary least squares (dWOLS) approach. Specifically, we utilize ChatGPT-5 and ScholarAI (Sept-Nov 2025 release) in the processes of identifying relevant sources for the literature review, creating summaries of papers, identifying gaps in research, and R code generation to implement methodology. Our findings show that current GenAI models lack the depth and contextual understanding required to accomplish these tasks without careful prompting and supervision of a knowledgeable researcher. Nonetheless, GenAI has potential to increase efficiency of tasks which take advantage of its search and summarization abilities, as well as basic code debugging and algorithm formation. We demonstrate that under a knowledgeable guide, GenAI can function as a research tool, but not as a substitute for methodological expertise.
Natalie Morosin, A. A. Nadi, M. Wallace· 0 citations
This systematic review examines recent progress in the pretraining and adaptation of LLMs for Low-Resource Languages (LRLs) and focuses on the ethics in AI practice, the development of corpora through communities, and interdisciplinary research collaboration among computational linguists, social scientists, and digital humanists.
Ismail Hossain, Mridul Banik, Fahmid Al Farid et al.· Computer Modeling in Enginee...· 0 citations
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
The broad adoption of Large Language Models (LLMs) has increased the need for human-curated datasets that serve as evaluation benchmarks. This need is particularly pronounced for non-English languages and for tasks that are inherently subjective and require multiple human perspectives. One such example is the development of benchmarks designed to assess the cultural awareness of LLMs. Statistics and data science courses offer a potential setting for developing such benchmarks while teaching students to apply LLM evaluation techniques using statistical inference. This paper presents a pilot project in which students in a statistics course within a data science engineering program created culturally diverse multiple-choice questions, generated answers using LLMs, and applied statistical methods to assess model accuracy. Student feedback indicated the project was engaging and useful for learning, while also highlighting a notable reliance on LLMs, particularly for interpreting statistical results. The resulting dataset comprises 1355 multiple-choice questions across 18 categories, including language, social media, and politics. After filtering valid items, the dataset was used to evaluate both closed- and open-source LLMs. Results show that the Gemini (closed-source) and Qwen (open-source) model families achieved the best performance, with improvements linked to model size, reasoning capabilities, and access to search tools. The best closed-source model achieved an accuracy of 97.66%, whereas the best open-source model achieved an accuracy of 79.07%. Qualitative analyses of errors in the filtering procedure and model reasoning process point to possible explanations into the challenges LLMs face when handling culturally specific content. Furthermore, results support a cultural injection hypothesis, whereby cultural knowledge is embedded during pretraining and accessed through instruction tuning. Through this work, we aim to demonstrate how statistics and data science courses can provide productive contexts for developing open-source benchmarks for non-English languages while also enriching students’ learning experiences. The dataset is publicly available.
Denis Iorga, Razvan Muntean, Mihai Masala et al.· Electronics· 0 citations
Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.
Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili et al.· 0 citations