Skip to content
Open access

From Medical Records to AI-Ready Datasets: A Practical Guide for Clinical Researchers

Aug 2026 · Journal of Clinical Medicine · Vol 15, pp. 6297 · 0 citations · 42 references
Medicine

TL;DR

A physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis to improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research is proposed.

Abstract

Background: Medical artificial intelligence (AI), machine learning (ML), and deep learning (DL) studies frequently begin with datasets collected for routine care rather than for computational modeling. Such datasets may contain inconsistent variables, heterogeneous measurement time points, unexplained NaN values, poorly defined outcomes, missing metadata, and insufficient documentation, which can compromise model development before any algorithm is selected. Methods: This Technical Note proposes a physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis. The guide was developed as a practical framework organized around pre-modeling decisions, including the clinical task, cohort, minimum common dataset, outcome definition, predictor variables, measurement timing, missing-data logic, standardization, non-tabular data linkage, data dictionary, and validation readiness. Results: The proposed guide translates AI-readiness principles into concrete data-collection rules for clinical, laboratory, imaging, physiological-signal, textual, follow-up, and multimodal data. It emphasizes clinically consistent data acquisition, reliable target labeling, explicit missing-data logic, patient-level linkage, structured metadata, and validation feasibility. A structured checklist and scoring approach are also proposed as practical pre-modeling assessment tools to classify datasets as not ready, exploratory only, ML-ready with limitations, or AI-ready for model development. Conclusions: Medical AI-readiness should be established before model development begins. By helping physicians collect, structure, and document data more consistently, the proposed guide may improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research.

Read PDF

Similar papers

Review Open access Mar 2026

Artificial Intelligence in Healthcare Practice: Validation, Fairness, and Regulatory Challenges: A Systematic Review

AI demonstrates strong potential to improve the effectiveness, safety, and quality of healthcare, however, broader clinical adoption remains constrained by regulatory requirements, interpretability gaps, data quality issues, and workflow integration challenges, underscoring the need for stronger validation practices and more implementation-focused research.

Ghulam Hussain Noori, Shaista Bibi, Seung Won Lee · 0 citations
Book Open access Aug 2026

OneEHR: Reproducible and AI Agent-Ready Longitudinal EHR Analysis Toolkit

Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, temporal preprocessing, label definitions, patient-level splits, and evaluation protocols can overshadow the methods being compared, making fair comparison and model selection difficult in practice. This tutorial presents OneEHR, an open-source toolkit that defines a unified experiment contract for modern EHR modeling and enables head-to-head comparison among conventional, neural, LLM-based, and agentic methods through a single configuration-driven interface. The three-hour hands-on session interleaves a methodological survey with guided practice: participants will learn why EHR experiments are vulnerable to leakage, distribution shift, and irreproducible preprocessing, and then use OneEHR to configure, execute, compare, and interpret experiments across this method spectrum. Attendees will leave with reusable configurations and a practical framework for integrating reproducible workflows into their own clinical AI research. Code and documentation are available at https://medx-pku.github.io/OneEHR/.

Yinghao Zhu, Zixiang Wang, Lei Gu et al. · 0 citations
Review Open access Aug 2026

Deep learning and generative AI for medical imaging and clinical decision support systems: a structured critical review

It is argued that LLM-based CDSS are not supported for routine autonomous use and require clinician supervision, and set out a research agenda centred on validation, governance and human-in-the-loop deployment.

A. Babu, A. J. Nehemiah, V. Jagadeep et al. · 0 citations
Open access Sep 2026

Less Can Be Better: Decomposing Clinical Data Modalities in Large Language Model-based Healthcare Applications

Objective To systematically examine different clinical data modalities in large language models (LLMs) and multimodal large language models (MLLMs), and to quantify the contribution of data modalities in early inpatient risk prediction and decision support tasks. Materials and Methods We conducted a systematic analysis using MIMIC-IV, MIMIC-IV-Note, and MIMIC-CXR-JPG datasets to create a unified cohort of 22,254 hospital admissions containing structured electronic health records (EHRs), radiology reports (clinical notes), and chest X-ray images. We evaluated general-purpose and medical-adapted LLM/VLMs across uni-, bi-, and tri-modal configurations on two risk prediction tasks (in-hospital mortality and length-of-stay [LOS] prediction) and two clinical decision support (CDS) tasks (discharge diagnosis phenotyping and medication-use prediction). Results For risk prediction tasks, structured EHR data alone achieved the best or comparable performance (best mortality AUROC: 0.849; LOS AUROC: 0.868), with limited incremental benefit observed from adding radiology reports or medical images. For CDS tasks, multimodal integration yielded substantial improvements: the best tri-modal configuration achieved F1-scores of 0.589 (diagnosis) and 0.405 (medication), representing 21.4% and 18.4% improvement over the best unimodal approach. Radiology reports consistently outperformed raw single-view chest radiographs as a supplementary modality. MLLMs demonstrated better zero- and few-shot performance than unimodal LLMs. Multi-view imaging consistently improved performance over single-view across all tasks. Conclusion The benefits of multimodal data integration are task-dependent. Healthcare LLMs should examine clinical data modalities according to specific tasks for efficient integration. These findings provide practical guidance for designing efficient clinical decision support systems.

Cheng Peng, Mengxian Lyu, Ziyi Chen et al. · 0 citations
Open access Sep 2026

From Routine Healthcare Data to Research-Ready Cohorts: A Structured Preprocessing Framework.

Routine clinical data offer significant potential for data-driven research and real-world evidence. However, such data are typically heterogeneous, irregular, and not designed for research use. This has a detrimental effect on the transformation of the data into analysis-ready datasets. The following paper sets out a preprocessing framework that converts routine clinical data into research-ready datasets. The approach was developed in the context of a use case concerning kidney transplant rejection diagnostics, with the objective of reconstructing clinically meaningful treatment trajectories. The proposed framework is comprised of three distinct steps: firstly, recursive event classification (rEC), then context-sensitive event annotation (cEA), and finally, context-based time series harmonisation (cTSH). The purpose of the framework is to transform irregular measurements into standardised temporal representations. When applied to routine data, the framework generated structured, comparable datasets that resemble clinical study data. These enable statistical modelling, machine learning, and systematic data reuse.

Matthias Katzensteiner, Darian Liehr, Oliver J. Bott · 0 citations
Open access Aug 2026

Can GPT Be Used as an Alternative Prediction Model to Traditional Machine Learning and Neural Networks on Low-Volume Clinical Data?

The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.

S. Bin Akter, S. Akter, D. Eisenberg et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.