A Standard-Constrained Evaluation Framework for Engineering-Oriented Applications of Large Language Models With Retrieval Result Unitization
Abstract
Large language model (LLM) applications are rapidly moving from general capability demonstrations to domain-specific engineering applications. Unlike public benchmark evaluations, project acceptance testing emphasizes whether a delivered system can provide correct retrieval, faithful generation, traceable evidence, and operationally usable answers under fixed business data, knowledge-base boundaries, and delivery requirements. Existing evaluation practices are often insufficient for Text-to-SQL, retrieval-augmented generation (RAG), and GraphRAG systems because their retrieval outputs differ substantially in structure and are difficult to evaluate under a unified criterion. To address this problem, this paper proposes a standard-constrained evaluation framework that combines automatic assessment and expert scoring for engineering-oriented LLM applications. The framework converts heterogeneous retrieval outputs into answer semantic units (ASUs), aligns predicted ASU lists with gold ASU lists, and computes retrieval-stage precision, recall, and F1-score using a unified true-positive, false-positive, and false-negative protocol. For the generation stage, ROUGE-L is adopted to measure the sequence-level similarity between system answers and reference answers, while domain-expert scoring is used to complement automatic metrics in terms of business correctness, completeness, traceability, and usability. A prototype experiment on a desensitized government intelligent question-answering system demonstrates that the proposed framework can reduce repetitive assessment cost while maintaining high agreement with expert judgment. The framework provides a reusable, interpretable, and auditable basis for acceptance testing, model selection, and iterative optimization of engineering-oriented LLM applications.