Skip to content

Author

Katharina von der Wense

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Calibration as a First-Class Criterion in LLM Evaluation

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without...

Mario Sanz-Guerrero, Katharina von der Wense · 0 citations

Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions

Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cult...

M. Bui, Mario Sanz-Guerrero, Abteen Ebrahimi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.