Skip to content
Open access

Domain-Specific Retrieval-Augmented Generation for Metallurgical R&D Knowledge Bases: A Hybrid Graph-Enhanced Approach

Aug 2026 · International Conference on Data Technologies and Applications · Vol 11, pp. 216 · 0 citations · 24 references

TL;DR

This work evaluates a confidence-adaptive graph-enhanced retrieval layer for metallurgical RAG in a controlled synthetic benchmark with explicitly specified generation and evaluation rules and shows how the retrieval rule behaves under controlled conditions.

Abstract

Metallurgical R&D search is difficult for a practical reason: useful evidence is rarely defined by one keyword. A production-support question can depend at the same time on material grade, process route, defect mechanism, property, test method, and numerical conditions. Conventional retrieval-augmented generation (RAG) pipelines largely treat document chunks as independent text and can therefore miss relations that matter for process monitoring, fault diagnosis, and engineering decision support. We evaluate a confidence-adaptive graph-enhanced retrieval layer for metallurgical RAG in a controlled synthetic benchmark with explicitly specified generation and evaluation rules. The benchmark contains 300 generated heterogeneous records derived from a seven-block source distribution and 30 material–process–defect–property archetypes, together with 180 frozen queries: 60 exact, 60 paraphrased, and 60 multi-hop. The main run evaluates robustness to incomplete structured metadata, with 10% missing and 4% erroneous categorical fields. Entity and relation extraction from raw documents is outside the evaluated scope. We compare BM25, TF-IDF, latent semantic analysis, a sparse + dense hybrid, graph-only retrieval, two ablations, and the proposed adaptive hybrid. On the complete query set, the proposed method obtains MRR = 0.992, Precision@5 = 0.980, Recall@10 = 0.859, and nDCG@10 = 0.948. Relative to the sparse + dense hybrid, nDCG@10 increases by 0.186 (24.4%); the paired 95% bootstrap interval is in the range of 0.166–0.206, and the Holm-adjusted Wilcoxon p-value is 2.59 × 10−29. Under severe degradation with 40% missing and 16% erroneous metadata, the adaptive method retains mean nDCG@10 = 0.791, compared with 0.650 for graph-only retrieval and 0.762 for the metadata-independent sparse + dense hybrid. A 5000-run Monte Carlo analysis estimates 6763 chunks and 58.70 MB for indexed vectors plus metadata at a 512-token chunk size and 64-token overlap. The results show how the retrieval rule behaves under controlled conditions; they are not evidence of plant-level effectiveness or of the quality of generated answers. Those questions require external, expert-labeled validation.

Read PDF

Similar papers

Open access Jul 2026

Agnostic Multi-Source Retrieval-Augmented Generation for Documents and Database Question Answering

This study develops a multi-source Retrieval-Augmented Generation (RAG) based Question Answering (QA) system that automatically integrates heterogeneous knowledge sources through a unified source parameter to enhance knowledge transfer and question answering for organizational support and employee onboarding.

Krisna Dwi Setya Adi, Ivan Michael Siregar · 0 citations
#small language model Review Open access Sep 2026

Evaluating LLM-Based Retrieval-Augmented Generation for Soil Science Question Answering

Retrieval-augmented generation (RAG) systems for scientific literature require evidence-based choices of document segmentation, representation, retrieval, and generation components, particularly when the source collection varies in topical specificity and document structure. This study addresses the lack of an end-to-e...

Karla Topić, Marina Bagić Babac, Vedran Mornar · 0 citations
Aug 2026

A large language model-based question-answering system for crack information

This research provides a highly accurate, scalable, and reliable framework for automated bridge defect analysis, offering a practical methodology to enhance data utilization in bridge management.

Lu-yang Zhang, Xuzhao Lu, Fengzong Gong et al. · 0 citations
Jul 2026

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.

Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al. · 1 citation
Open access Aug 2026

Towards Building a Multi-Source Heterogeneous Knowledge Graph for Complex Material Question Answering

Results indicate that integrating multi-source domain knowledge with relation-preserved retrieval and attribute-supported filtering provides more focused and inspectable evidence, thereby supporting more accurate complex material question answering.

Peize Li, Xi Guo, Nan Yin et al. · 0 citations
Book Open access Aug 2026

AutoSchema: Self-Prompted Schema Induction and Evidence-Grounded Extraction for Materials Science Literature

Scientific papers in materials science contain critical experimental details (e.g., reagents, synthesis conditions, and measured properties) for building structured databases and enabling downstream analysis, but large-scale structured extraction remains difficult. Classic rule-based systems rely on hand-written patter...

Mingfang Zhu, Yixin Chen, Zhiling Zheng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.