Skip to content
Review Open access

LLM-driven materials knowledge extraction: multimodal parsing, ontology, and agentic systems

Jul 2026 · Journal of Materials Informatics · Vol 6 · 0 citations · 93 references

TL;DR

It is argued that reliability is the decisive criterion for large-scale deployment, and synthesize failure modes, layered defenses, and evaluation protocols that connect source grounding, ontology constraints, physical verification, and human-in theloop review.

Abstract

Extracting reliable knowledge from unstructured materials literature remains a central bottleneck for data-driven and AI-enabled materials discovery. Large language models (LLMs) are reshaping this task by integrating multimodal document parsing, ontology-guided semantic grounding, structured extraction, and agentic verification into increasingly unified workflows. This review analyzes these developments through a Perception–Cognition–Action lens. At the perception layer, we examine how scientific document parsers, multimodal LLMs, table and chart readers, and optical chemical-structure-recognition systems convert visually rich papers into computable evidence. At the cognition layer, we discuss how ontologies and knowledge graphs constrain LLM outputs, support entity alignment, and reduce semantic ambiguity. At the action layer, we compare schema-based extraction, schema-free discovery, and agentic extraction as a control–coverage–autonomy spectrum rather than a simple succession of tools. We further argue that reliability is the decisive criterion for large-scale deployment, and synthesize failure modes, layered defenses, and evaluation protocols that connect source grounding, ontology constraints, physical verification, and human-in-the-loop review. By distinguishing demonstrated extraction capabilities from more speculative AI-scientist and self-driving-laboratory visions, this review provides a comparative and risk-aware account of how LLM-driven systems can produce evidence-linked, physically meaningful, and reusable materials knowledge.

Read PDF

Similar papers

Open access Aug 2026

A Synergistic Knowledge Graph and LLM-Driven Framework for Intelligent Process Decision-Making Systems

The proposed knowledge graph construction method for the workpiece machining distortion domain is proposed, together with an intelligent decision-making framework driven by the collaboration of knowledge graphs and large language models, providing a feasible pathway for the structured organization, intelligent retrieva...

Deguo Yao, Zhaoze Sun, Jie Gao et al. · 0 citations
Preprint Aug 2026

Toward Effective and Reliable LLM Agents via Dynamic Ontology

OaK is presented, an ontology-as-a-kernel framework that dynamically constructs and refines task-oriented ontologies for LLM agents and shows that OaK improves standard LLM agents, strengthens evidence grounding, and boosts the reliability of multi-step reasoning.

Xiaohui Zhang, Ze-Qun Sun, Cheng Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs

This paper presents a self-demonstration-driven approach that combines a neuro-symbolic task decomposition with a novel mechanism for automatically generating pattern-guided, dependency-aware demonstrations to address the integration challenge of heterogeneous relational databases into a centralized ontology.

Siddhesh Thombre, Manasi S. Patwardhan, Sunita Sarawagi · 0 citations
Conference Open access Aug 2026

AI-Driven Knowledge Externalisation: From Unstructured Documents to Structured Data Models

The findings suggest that AI-based structured extraction may redefine how organisations formalise expertise, shifting from document-centric storage toward schema-driven knowledge architectures.

Dilyan Georgiev, E. Gourova · 0 citations
Review Open access Aug 2026

Large language models for knowledge-centric scientific intelligence: methods, challenges, and lessons from geoscience

This review critically examines the emerging literature on geoscience-oriented LLMs (GeoLLMs), focusing on the tasks, construction strategies, evaluation needs, and unresolved challenges that distinguish them from generic LLM applications.

Jian-Hua Ma, Yong-Zhang Zhou, Luhao He et al. · 0 citations
Conference Aug 2026

Ontology-Guided Pedagogically Meaningful Knowledge Component Extraction from Code

Accurate extraction of Knowledge Components (KCs) is critical for fine-grained learner modeling in programming education. Yet existing approaches remain limited: manual Q-matrices ignore solution variability; Abstract Syntax Tree (AST) based methods may not produce pedagogically meaningful KCs; and Large Language Model...

Mathangi Krishnathasan, K. Hewagamage, E. Hettiarachchi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.