Aug 2026· Proceedings of the 2026 ACM Symposium on Document Engineering· pp. 1-4· 0 citations· 13 references
TL;DR
This paper presents an end-to-end framework for automatic instrument-tag extraction and association from scanned P&IDs, and demonstrates accurate tag reconstruction, highly reliable symbol-tag associations, and a substantial reduction in manual annotation effort, providing an effective foundation for large-scale P&ID digitization.
Abstract
Piping and Instrumentation Diagrams (P&IDs) are essential engineering documents, but many remain available only as scanned PDFs, limiting their integration into digital workflows. Automatic extraction of structured information from these drawings is challenging due to large document sizes, small text annotations, and ambiguous symbol-tag relationships. This paper presents an end-to-end framework for automatic instrument-tag extraction and association from scanned P&IDs. A tiled OCR strategy with graph-based text merging improves text completeness and reconstructs fragmented engineering tags. For symbol detection, an RF-DETR model fine-tuned on the Dataset-P&ID benchmark achieves 99.96% mAP@50 and 99.97% precision, while SAHI-based inference slicing improves performance on large drawings. To automate symbol-tag association, we formulate the problem as a minimum-cost bipartite matching task that combines geometric, semantic, and spatial cues and solves it globally using the Hungarian algorithm. Results on Dataset-P&ID demonstrate accurate tag reconstruction, highly reliable symbol-tag associations, and a substantial reduction in manual annotation effort, providing an effective foundation for large-scale P&ID digitization. The code is available at https://github.com/dimitri009/STA.
The problem of parsing heterogeneous PDF documents is still unresolved because of the use of complex layouts, embedded watermarks, and the disposition of the current systems to generate incoherent knowledge graphs (KGs) consisting of hundreds of nodes that are loosely connected and do not contribute a lot to education....
ReG-TG is proposed, a retrieval-augmented framework for repository tag recommendation that integrates hybrid retrieval, a tag co-occurrence knowledge graph, and Chain-of-Thought reasoning within a large language model (LLM)-based architecture and consistently outperforms representative baselines.
Min Wang, Shan-Shan Wu, Wanjia Lv et al.· Applied Sciences· 0 citations
An engineering format-aware masking strategy is proposed that improves entity recognition for phase-related structures, engineering abbreviations and equipment hierarchy fragments and achieves precision, recall and F1 scores of 0.90, respectively.
Ya Wu, Cui-Ru Yang, Yao Yao et al.· Energies· 0 citations
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ p...
Christoph Walser, Mauricio Fadel Argerich, Jonathan Furst· 0 citations
Automatic routing is a necessary first stage in many scanned-document workflows, yet small systems often face an unattractive choice between brittle keyword rules and computationally expensive deep classifiers. Whole-page rules discard word placement, while a forced prediction can silently misroute an ambiguous page. T...
Zhen Wang, Zheng Yong, Yao Ma et al.· 2026 3rd International Confe...· 0 citations
SmolDocling, a compact 256M-parameter vision-language model (VLM), is fine-tune to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing.
A. Gurbuz, A. Nassar, Christoph Auer et al.· IEEE International Conferenc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.