Skip to content
Open access

A Multilevel Visual and Textual Framework for Near-Duplicate Diagram Detection in Electronic Documents

Sep 2026 · Information · Vol 17, pp. 897 · 0 citations

TL;DR

A cascaded multimodal framework combining perceptual hashing, Siamese Residual Network with 18 layers, Siamese Residual Network with 18 layers, Distillation with No Labels, ver. 2 (DINOv2)visual representations, and a text-similarity classifier is proposed to detect near-duplicate diagram detection in electronic documents.

Abstract

Near-duplicate diagram detection in electronic documents is challenging because diagram identity depends on graphical structure, spatial composition, and textual labels, while reused images may undergo compression, cropping, rotation, photometric changes, or perspective distortion. This study proposes a cascaded multimodal framework combining perceptual hashing, Siamese Residual Network with 18 layers (Siamese ResNet18), Distillation with No Labels, ver. 2 (DINOv2)visual representations, and a text-similarity classifier. A controlled benchmark was constructed from Artificial Intelligence 2D Diagram Dataset (AI2D) using Light, Medium, and Hard transformations, with source-grouped splitting by base_id to prevent leakage across training, validation, and test sets. Perceptual hashing achieved Area Under the Receiver Operating Characteristic Curve (ROC-AUC) = 0.7483, while Siamese ResNet18 increased ROC-AUC to 0.8599. DINOv2 provided the strongest visual performance, achieving Accuracy = 0.9933, F1-score = 0.9933, ROC-AUC = 0.9992, and Average Precision = 0.9994; F1-score remained 0.9901 for Hard transformations. Visual fusion increased ROC-AUC to 0.9995, and full multimodal fusion reached ROC-AUC = 0.9998. At an early-exit threshold of 0.95, 27.6% of pairs were resolved at the hashing level. These results support the coarse-to-fine design on the constructed AI2D-derived benchmark. A targeted hard-negative stress test revealed substantially higher false-positive rates under deliberately matched spatial layouts, with an overall False Positive Rate (FPR) of 0.48 for DINOv2 and 0.16 for full multimodal fusion. Generalization to naturally reused or redrawn diagrams, larger and more diverse hard-negative collections, and Optical Character Recognition (OCR)-derived text remains to be evaluated.

Read PDF

Similar papers

Open access 2026

A Confidence-Aware Hybrid OCR and Visual Retrieval Framework for Arabic Document Images

The paper will suggest and mathematically model a confidence-aware hybrid retrieval system in Arabic document images with the use of keyword queries, and presents Arabic-oriented framework, reproducible evaluation protocol, measured auxiliary evidence and the identification of the dedicated Arabic retrieval benchmark t...

A. S. Ibrahim · 0 citations
#small language model Open access Aug 2026

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

SmolDocling, a compact 256M-parameter vision-language model (VLM), is fine-tune to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing.

A. Gurbuz, A. Nassar, Christoph Auer et al. · 0 citations
#small language model Open access Sep 2026

Character-Level Visual Guidance for Arbitrarily Shaped Scene Text Detection and Recognition

Study of explicit character-level guidance for scene text detection and recognition indicates that character semantics for detection and local visual evidence for decoding are an effective way to improve robustness in difficult scene text.

Li-Jia Chen, Hu Lin, Dan Chen · 0 citations
Aug 2026

Chinese Painting Semantic Resource for Resolution-Aware Digitisation and Structured Heritage Access

Chinese painting digitisation poses a distinctive heritage-informatics challenge: culturally salient visual structures such as brushwork, intentional blank space, and text–image coexistence are difficult to document, access, and research with general-purpose collection and computational representations. We present the...

Hao-Rui Yu, Jiao Xu, Ting-Ting Yang et al. · 0 citations
Preprint Aug 2026

AcroMELD: Recovering Interactive PDF Forms with Structure-Aware Graph Set Transformers

Interactive PDF form fields are often absent from documents that visually resemble forms, leaving users unable to enter data without printing or external editing tools. Detecting the missing widgets is difficult because a field may be indicated by several overlapping cues, born-digital PDFs expose useful but incomplete...

Samuel Abramov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.