Skip to content
Book Open access

Towards Automated P&ID Digitization: Graph-Based OCR Consolidation and Global Symbol-Tag Association

Aug 2026 · Proceedings of the 2026 ACM Symposium on Document Engineering · pp. 1-4 · 0 citations · 13 references

TL;DR

This paper presents an end-to-end framework for automatic instrument-tag extraction and association from scanned P&IDs, and demonstrates accurate tag reconstruction, highly reliable symbol-tag associations, and a substantial reduction in manual annotation effort, providing an effective foundation for large-scale P&ID digitization.

Abstract

Piping and Instrumentation Diagrams (P&IDs) are essential engineering documents, but many remain available only as scanned PDFs, limiting their integration into digital workflows. Automatic extraction of structured information from these drawings is challenging due to large document sizes, small text annotations, and ambiguous symbol-tag relationships. This paper presents an end-to-end framework for automatic instrument-tag extraction and association from scanned P&IDs. A tiled OCR strategy with graph-based text merging improves text completeness and reconstructs fragmented engineering tags. For symbol detection, an RF-DETR model fine-tuned on the Dataset-P&ID benchmark achieves 99.96% mAP@50 and 99.97% precision, while SAHI-based inference slicing improves performance on large drawings. To automate symbol-tag association, we formulate the problem as a minimum-cost bipartite matching task that combines geometric, semantic, and spatial cues and solves it globally using the Hungarian algorithm. Results on Dataset-P&ID demonstrate accurate tag reconstruction, highly reliable symbol-tag associations, and a substantial reduction in manual annotation effort, providing an effective foundation for large-scale P&ID digitization. The code is available at https://github.com/dimitri009/STA.

Read PDF

Similar papers

Conference Aug 2026

Layout-Aware Document Parsing and Knowledge Graph Extraction

The problem of parsing heterogeneous PDF documents is still unresolved because of the use of complex layouts, embedded watermarks, and the disposition of the current systems to generate incoherent knowledge graphs (KGs) consisting of hundreds of nodes that are loosely connected and do not contribute a lot to education....

Jayesh Jain, Yash Vipul Sinojia, Soshya Joshi · 0 citations
Open access Aug 2026

A Software Repository Tag Method Based on Hybrid Search and Graph Enhancement

ReG-TG is proposed, a retrieval-augmented framework for repository tag recommendation that integrates hybrid retrieval, a tag co-occurrence knowledge graph, and Chain-of-Thought reasoning within a large language model (LLM)-based architecture and consistently outperforms representative baselines.

Min Wang, Shan-Shan Wu, Wanjia Lv et al. · 0 citations
Open access Aug 2026

A Named Entity Recognition Method for GIS Defect Texts Incorporating an Engineering Format-Aware Masking Strategy

An engineering format-aware masking strategy is proposed that improves entity recognition for phase-related structures, engineering abbreviations and equipment hierarchy fragments and achieves precision, recall and F1 scores of 0.90, respectively.

Ya Wu, Cui-Ru Yang, Yao Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ p...

Christoph Walser, Mauricio Fadel Argerich, Jonathan Furst · 0 citations
Conference Aug 2026

A Lightweight Spatial-Anchor-Aware Confidence Routing Method for Heterogeneous Scanned Documents

Automatic routing is a necessary first stage in many scanned-document workflows, yet small systems often face an unattractive choice between brittle keyword rules and computationally expensive deep classifiers. Whole-page rules discard word placement, while a forced prediction can silently misroute an ambiguous page. T...

Zhen Wang, Zheng Yong, Yao Ma et al. · 0 citations
#small language model Open access Aug 2026

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

SmolDocling, a compact 256M-parameter vision-language model (VLM), is fine-tune to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing.

A. Gurbuz, A. Nassar, Christoph Auer et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.