Preprint
Jul 2026
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework.
Jeonghwan Choi, Taewon Yun, Minjeong Ban et al.
· 1 citation