Skip to content

Author

A. Kulinkina

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

Introduction: Large language models (LLMs) have the potential to strengthen clinical decision-making in low-resource primary healthcare (PHC) settings. However, most LLMs are developed and benchmarked in high-resource settings and evidence on their safety and contextual appropriateness in Sub-Saharan Africa remains limited. The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania. It aims to validate a pool of LLMs through clinical review of expert-generated vignettes. Methods and analysis: This is a fully crossed repeated-measures comparative evaluation study. In each country, experienced clinicians develop 200-250 hypothetical clinical vignettes reflecting realistic patient presentations and independently produce a human benchmark care plan for each. Vignettes are used to prompt a selection of six open-source and proprietary LLMs selected based on code availability, local hostability, and model size. During in-person workshops (valiDATAthons), independent clinical experts rate LLM- and human-generated responses in source-attribution masked side-by-side comparisons across five dimensions (clinical soundness, safety, contextual fit, clarity & completeness, and appropriate confidence). The primary endpoints are each LLM's overall performance profile and non-inferior safety profile, as compared to the human benchmark. At minimum, 358 evaluations per LLM (or 1,253 paired evaluations in total) are required per country. Ethics and dissemination: The study is approved by the EPFL Human Ethics Research Committee in Switzerland, Harvard T.H. Chan School of Public Health in the USA, KNH-UoN Ethics and Research Committee in Kenya, MUBAS Research Ethics Committee in Malawi, and MUHAS Research and Ethics Committee and National Institute for Medical Research in Tanzania. Findings will be reported according to the TRIPOD-LLM framework and shared with national ministries of health, disseminated at conferences and in peer-reviewed journals, and de-identified benchmark data will be released under FAIR principles.

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations
Open access Jul 2026

Federated modular clinical decision support networks for collaborative learning in resource-limited settings

Imperfect interoperability (IIO), where health facilities record different, often sparse subsets of clinical variables, remains a major barrier to deploying models trained with Federated Learning (FL) in global health settings. We introduce FedMoDN, a novel federated modular neural network architecture for collaborative learning across all features of an IIO distributed dataset, allowing healthcare facilities to use the full complement of their features without sharing, discarding, or imputing any data. We evaluate FedMoDN on a multi-site pediatric dataset comprising ~130,000 medical visits across 92 healthcare facilities in Tanzania and Rwanda. Across both internal and external validation health facilities, FedMoDN matches or surpasses models trained with centralized data sharing and competitive monolithic FL baselines, achieving a mean AUPRC of 0.80 versus 0.77 for the monolithic FL model on 18 external validation health facilities. Its relative advantage over a monolithic FL model rose from 4% (complete data) to 22% when 70% of test-time features were missing, and, unlike monolithic FL models, performance remained stable when health facilities contributed disjoint feature or label subsets. Furthermore, step-wise predictions provide clinically interpretable feature-attribution scores. By coupling IIO resilience with built-in interpretability, FedMoDN offers a promising decision support tool for resource-limited facilities sidelined by conventional FL.

Cécile Trottet, Jonathan Doenz, P. M. Mastel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.