Skip to content
Open access

LGF-Net: a spatial-frequency deepfake detection network with learnable gabor filters and gated cross-modal Fusion

Sep 2026 · Discover Computing · Vol 29 · 0 citations · 94 references

Abstract

Deepfakes have become increasingly realistic due to recent advances in face manipulation techniques, making reliable detection in unconstrained environments more challenging. Existing spatial-frequency deepfake detection methods often rely on fixed hand-crafted frequency transforms and simple fusion strategies, which may limit their adaptability and cross-dataset generalization. To address these limitations, we propose LGF-Net, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection. The Frequency Representation Module employs learnable Gabor filters and a frequency-aware attention mechanism to capture manipulation-specific spectral patterns. Moreover, the Spatial Representation Module uses multi-rate dilated convolutions to model both subtle local artifacts and long-range structural inconsistencies, while a gated cross-modal fusion module integrates the two representations into a compact forensic descriptor. Experimental results on FF++ (HQ), Celeb-DF (V2), DPDC, and DFD show that LGF-Net achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.