Skip to content
Conference

A Deep Learning Framework for Fraud Detection on Multimodal Financial Report Data

Aug 2026 · 2026 6th International Conference on Computer Science and Blockchain (CCSB) · pp. 387-390 · 0 citations · 12 references

Abstract

Financial statement fraud (FSF) poses severe threats to investor confidence and capital market stability, yet most existing detection models rely solely on a single modality, such as either financial ratios extracted from tables or textual disclosures in management discussion. Such single-view models fail to capture the cross-modal inconsistencies that often underlie fraudulent reporting. To address this limitation, we propose a Deep Learning Framework for Fraud Detection on Multimodal Financial Report Data (FDMFR) that integrates textual disclosures, tabular financial ratios, and visual layout cues into a unified end-to-end pipeline. The proposed framework consists of five stages: (1) modality-specific preprocessing; (2) modality-specific feature encoders (BERT+BiLSTM for text, MLP for tabular data, and CNN for visual layout); (3) attention-based cross-modal fusion that adaptively reweights the contribution of each modality; (4) a binary classification head producing a fraud probability; and (5) an interpretability output highlighting suspicious sentences and abnormal financial indicators. Extensive experiments on a private benchmark of real-world annual financial reports demonstrate that FDMFR achieves superior detection performance, with an F1-score of 0.892 and an AUC of 0.947, outperforming state-of-the-art single-modal and early-fusion baselines by 5.1–15.7 percentage points in F1-score. Ablation studies confirm that each modality and the attention-based fusion contribute complementary evidence, validating the necessity of joint multimodal reasoning for financial fraud detection.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.