Findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.
Sun-Yang Fu, M. J. Kwak, J. Ahn et al.· npj Health Systems· 0 citations
This work introduces MedJudge, a multimodal medical reward modeling method that supports interpretable, evidence-grounded, and clinically-aligned decision evaluation, and proposes UMLS-based Concept Overlap (UCO) to evaluate explanation quality, measuring concept-level alignment with clinician expectations.
Yun-Hong He, Kai Zhang, Jiarong Qian et al.· Proceedings of the 32nd ACM...· 0 citations
The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffering from two critical limitations: (1) data contamination, where test sets inadvertently leak into training corpora, leading to inflated per...
Zhiling Yan, D. Song, Zhe Fang et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces MedJudge, a multimodal medical reward modeling method that supports interpretable, evidence-grounded, and clinically-aligned decision evaluation, and proposes UMLS-based Concept Overlap (UCO) to evaluate explanation quality, measuring concept-level alignment with clinician expectations.
Yunhong He, Kai Zhang, Jiarong Qian et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.