Skip to content
Review Open access

Text2FHIRwallet: Automated Generation of FHIR Patient Summaries from Unstructured Cardiology Reports Using Fine-Tuned Portuguese Language Models—Development and Evaluation of a Health Professional Wallet

Aug 2026 · Applied Sciences · Vol 16, pp. 7795 · 0 citations · 24 references

TL;DR

Text2FHIRwallet demonstrates that domain-specific fine-tuning of a Portuguese-language pretrained language model achieves near-ceiling clinical NER accuracy, enabling scalable, interoperable, and privacy-compliant patient summary generation from unstructured cardiology text, offering an end-to-end pathway for integrating AI-driven NLP into clinical workflows and FHIR-based health information ecosystems.

Abstract

Background and Objectives: Cardiology departments generate large volumes of unstructured free-text reports that impose substantial manual review burdens on clinicians; at Hospital de Santa Maria—Portugal’s largest public hospital—manual review of 12,651 reports took approximately seven minutes per report, representing over 1475 h of avoidable administrative work. This study presents Text2FHIRwallet, a health professional digital wallet that automates extraction and structuring of clinical entities from unstructured Portuguese cardiology reports using fine-tuned Named Entity Recognition (NER) models and maps the results to Fast Healthcare Interoperability Resources (FHIR) R4 patient summaries. Materials and Methods: Following the Design Science Research Methodology (DSRM) and CRISP-DM, we fine-tuned four transformer-based models—BERTimbau Base, BERTimbau Large, Albertina PT-PT, and MediAlbertina—on 305 manually annotated cardiology reports (77,309 tokens; κ = 0.85 inter-annotator agreement) covering eight clinical entity types, drawn from a corpus of 12,651 anonymised documents. Entities were mapped to FHIR R4 resources and delivered through a secure, role-based mobile wallet (React Native). Evaluation comprised token-level NER benchmarking with bootstrapped confidence intervals and McNemar’s testing, FHIR mapping accuracy assessment on 100 manually reviewed reports, processing-efficiency measurement, and a usability pilot with 10 cardiologists (SUS, NPS). Results: MediAlbertina achieved the highest NER performance (macro F1 = 0.985, 95% CI: 0.979–0.990), significantly outperforming all baseline models (p < 0.01, McNemar’s test) and comparing favourably with—though not directly comparable to, given differing languages and datasets—published benchmarks such as GPT-4 (F1 = 0.962 in ophthalmology NER) and fine-tuned BERT models for lung cancer NER (F1 ≈ 0.85–0.90). FHIR mapping accuracy was 98% on 100 independently reviewed reports. Report processing time was reduced from approximately seven minutes to 15–30 s (93–96% reduction), with peak batch-inference throughput of up to 1000 reports/h under parallelised GPU load (observed end-to-end throughput in pilot deployment was approximately 250 reports/h). The pilot usability evaluation yielded a SUS score of 87 (excellent) and an NPS of 80. Conclusions: Text2FHIRwallet demonstrates that domain-specific fine-tuning of a Portuguese-language pretrained language model achieves near-ceiling clinical NER accuracy, enabling scalable, interoperable, and privacy-compliant patient summary generation from unstructured cardiology text, offering an end-to-end pathway for integrating AI-driven NLP into clinical workflows and FHIR-based health information ecosystems, with implications for administrative efficiency, care coordination, and clinical research in non-English-language settings.

Read PDF

Similar papers

Review Aug 2026

From text to insight: a systematic literature review of keyword and keyphrase extraction techniques for healthcare text mining

This review provides researchers and practitioners with a structured framework for method selection based on their specific constraints and identifies six prioritized research directions for future investigation, identifying critical research gaps including the preservation of multi-word clinical concepts, scarce evalu...

Mouhamed Gaith Ayadi · 0 citations
Conference Jul 2026

Generating Reliable Synthetic Clinical Discharge Summaries for Medical Text Analysis

Clinical text is an important part of healthcare systems because it is used to store and manage patient information in documents such as discharge summaries, doctor notes, and diagnostic reports. Among these documents, discharge summaries are especially important because they provide a brief overview of a patient’s dia...

Mohammad Imran, M. Irfan, Sajida Sultana.Sk et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Review Open access Aug 2026

NERFlow: A Workflow-Based Subsystem of FIT4NER for LLM-Assisted Medical Named Entity Recognition

Preparing training data for domain-specific medical Named Entity Recognition (NER) involves a trade-off between annotation quality, expert effort, and data privacy: manual annotation is costly, whereas cloud-based Large Language Models (LLMs) raise concerns about the control of sensitive clinical text. This article int...

Florian Freund, Philippe Tamla, Bao Tran et al. · 0 citations
Open access Feb 2025

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

This study evaluated both closed-source and open-source large language models for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation and found that model performance could deviate largely based on prompting style...

W. Jang, Sharmin Sultana, Zonghai Yao et al. · 1 citation
Open access Jul 2026

Benchmarking large language models for clinical data extraction from Portuguese medical notes in a university hospital

The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.

Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.