A Graph-Enhanced Multimodal Transformer Model for Fine-Grained Document Parsing
The digitization of documents across industries has created an urgent need for intelligent systems capable of extracting structured information from unstructured layouts. While transformer-based models like LayoutLMv3 have advanced document understanding, they struggle to capture relational dependencies between spatial...