Open access
2026
Evaluating LLM-as-a-Judge for Medical Term Simplification
The reliability and robustness of LaaJ for specialized medical knowledge is investigated by evaluating six LLMs for their judgment capabilities on three dimensions: correctness, readability, and completeness, and it is observed that hallucinations in LaaJ setups can be mitigated by epistemic markers.
Ioana Buhnila, Aman Sinha, Rohit Agarwal et al.
· BioNLP@ACL · 0 citations