Vi-FusionQA: A Unified Benchmark and Baseline Study for Vietnamese Question Answering
Abstract
Vietnamese question answering (QA) is increasingly important for customer support and digital service applications, yet existing Vietnamese QA resources remain limited in scale, domain coverage, and task format. In this paper, we introduce Vi-FusionQA, a public unified benchmark for cross-domain Vietnamese question answering. Rather than targeting a single deployment setting, Vi-FusionQA is designed as a general-purpose benchmark, with OTT customer support used as a motivating example of a practical QA application. Vi-FusionQA aggregates multiple Vietnamese and multilingual QA resources into a shared context-question-answer schema, resulting in over 100,000 QA pairs with source-level traceability. We further enrich the dataset with LLM-generated pseudo reasoning types and difficulty labels to support exploratory dataset analysis. To establish baseline performance, we fine-tune representative open-weight LLMs, including Qwen and Llama variants, using LoRA and evaluate them with standard QA metrics. Experimental results show that Qwen3-4B-Instruct-2507 achieves the best full-dataset performance, reaching 68.22% F1 and 37% Exact Match. We also report system-level efficiency metrics, including training speed, GPU utilization, and power consumption, to support practical model selection for Vietnamese QA applications.