Abstract concepts such as harmfulness, legality, privacy, and fraud present a major challenge for multimodal large language models (MLLMs), as they require reasoning across visual, linguistic, and contextual information rather than direct perception. This paper investigates three questions: whether supervised alignment...
Ying-Xu Wang, Oliver Lemon· Companion Publication of the...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.