Traffic Law Question Answering over dashcam videos requires both accurate visual evidence selection and reliable multimodal reasoning. This task is especially challenging in Vietnamese traffic scenes, where road environments are dense, diverse, and highly dynamic. In this paper, we present a resource-efficient two-stag...
Van-Hoang Le, Duc-Vu Nguyen, Kiet Van Nguyen et al.· International Conference on...· 0 citations
Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and...
T. Dang, H. Nguyen, Kiet Van Nguyen· International Conference on...· 0 citations
Paraphrase identification remains challenging when sentence pairs exhibit high lexical overlap but subtle semantic differences, as models often rely on surface similarity rather than true meaning. Existing benchmarks such as PAWS highlight this issue, but comparable resources for Vietnamese are still lacking. In this p...
Sang Quang Nguyen, Kiet Van Nguyen, N. L. Nguyen et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.