Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared through clinician-oriented online resources. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical reasoning and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4 percent. It also generated more clinically coherent responses on the ThoughtMed-1M test set, outperforming state-of-the-art models by 3 to 5 percent across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al.· 0 citations
Background Lung adenocarcinoma (LUAD), which is the leading subtype of non-small cell lung cancer (NSCLC), poses considerable difficulties in accurate prognostic assessment and targeted therapeutic options. Cell proliferation-related genes (CPGs) and mitochondrial biogenesis-related genes (MBGs) play critical roles in tumor metabolic reprogramming; however, their prognostic value and molecular mechanisms in LUAD are poorly understood. This study aims to construct a CPG/MBG-based prognostic risk model for LUAD, evaluate its clinical utility in predicting prognosis and immunotherapy response, and experimentally validate the functional role of key model genes in LUAD progression. Methods By utilizing The Cancer Genome Atlas (TCGA)-LUAD and GSE72094 datasets, this investigation formulated a risk scoring model through differential expression screening combined with least absolute shrinkage and selection operator (LASSO)-Cox regression analysis. The molecular characteristics and clinical implications of the risk model were investigated via immune microenvironment evaluation, genomic alteration analysis, and drug sensitivity prediction. The functional contributions of key genes were further substantiated using quantitative reverse transcription polymerase chain reaction (qRT-PCR), commercial assay kits, the JC-1 fluorescent probe, the Cell Counting Kit-8 (CCK-8), Transwell invasion assays, and wound healing assays. Results A risk model based on seven CPGs and MBGs (PLK1, HMMR, CYP27A1, LDHA, NPAS2, KRT17, CIDEC) showed reliable predictive performance in both GSE72094 and the TCGA-LUAD cohorts. Enhanced tumor heterogeneity and an immunosuppressive microenvironment were observed in the high-risk group. Drug sensitivity analysis indicated that the risk model could guide personalized treatment strategies; for instance, high-risk patients showed increased susceptibility to agents such as docetaxel and 5-fluorouracil. In vitro experiments demonstrated that the key gene CIDEC exhibited upregulated expression in LUAD tissues and cells. Knockdown of CIDEC led to enhanced cellular energy metabolism and increased mitochondrial membrane potential, while also effectively suppressing cell invasion, proliferation, and migration. Conclusions The established MBGs/CPGs prognostic model provides a novel tool for stratified treatment planning in LUAD, underscoring the crucial roles of cellular proliferation and mitochondrial biogenesis in tumor progression. Functional validation of CIDEC offers experimental support for the development of potential therapeutic strategies.
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Ling-Xuan Hou, Yu-Hua Xie, Yue Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.