Skip to content
Preprint

MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

Jul 2026 · 0 citations · 42 references
Computer Science

TL;DR

MorphUNet is the first diffusion-based morphing framework using trainable parent-separated dual cross-attention inside the denoising U-Net: a Biometric Transport Layer carrying parent-specific identity evidence through denoising, attending to each parent separately before combining residuals via the morphing parameter alpha.

Abstract

Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We introduce MorphUNet, a diffusion morphing framework formulating two-parent generation as alpha-controlled biometric transport: each parent is decomposed into CLIP appearance and ArcFace identity evidence, aligned into a CLIP-compatible token space, with the two contributors preserved as separate identity-aware token banks. To our knowledge, MorphUNet is the first diffusion-based morphing framework using trainable parent-separated dual cross-attention inside the denoising U-Net: a Biometric Transport Layer carrying parent-specific identity evidence through denoising, attending to each parent separately before combining residuals via the morphing parameter alpha. DDIM-inverted latent interpolation gives a coherent denoising start, while weaker-parent-guided selection favours morphs maximising the lower parent-similarity score, reducing collapse toward one contributor. We evaluate MorphUNet against three state-of-the-art baselines (StableMorph, MIPGAN-II, and MorDIFF) on FEI and FRLL using six recognition systems, and propose CFD-based unseen-identity stress testing across gender and ethnicity pairing, demographic shifts, and parent-similarity extremes. MorphUNet achieves the best Morphing Attack Potential (MAP) when at least three of six systems are fooled by one morph, reaching 0.919 on FEI and 0.886 on FRLL, and obtains the best FID on both datasets (35.19 FEI, 44.86 FRLL). It also gives the highest APCER at 5% BPCER in the same-dataset setting, and remains highly difficult to detect under cross-dataset transfer, with APCER 0.996 on FEI and 0.946 on FRLL. The full evaluation analyses MAP, MAD, per-system vulnerability, identity balance, image quality, top/bottom-similarity stress tests, and CFD unseen-identity robustness.

View source

Similar papers

Preprint Aug 2026

Face Re-morphing: Differential Morphing Attack Detection via Feature-Space Similarity Changes

Face morphing attacks pose a serious threat to face recognition systems because a single morphed document image can be matched to multiple contributors. Differential morphing attack detection (D-MAD) addresses this threat by comparing a document image with a trusted live image, but existing methods often rely on static feature differences, constituent-face reconstruction, or multi-cue fusion. This paper proposes Face Re-morphing, a D-MAD method that uses the feature-space response to an additional morphing operation as a detection cue. Given a document image and a trusted live image, the proposed method generates a re-morphed image and uses the change between the document--live and live--re-morphed cosine similarities as the detection score. Experiments on FRLL-Morphs and FEI Morph show that the proposed cue is effective across different morphing conditions, re-morphing methods, and face recognition models. Comparisons with existing methods show favorable results on AMSL and indicate that the proposed method performs well under the Criminal condition on FEI Morph Version~1, particularly when using MorDIFF. These results indicate that re-morphing-induced similarity change provides a complementary cue for D-MAD.

Jie Jin, M. Nishigaki, Tetsushi Ohki · 0 citations
Preprint Jul 2026

Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models

Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is especially vital for security systems, biometric authentication, and privacy sensitive applications, where any drift in identity integrity can undermine trust and functionality. We introduce Diff-ID, a diffusion based framework that enforces identity consistency while delivering photorealistic quality. Central to our approach is a custom 210K image dataset synthesized from CelebA-HQ, FFHQ, and LAION-Face and captioned via a fine tuned BLIP model to bolster identity awareness during training. Diff-ID integrates ArcFace and CLIP embeddings through a dual cross attention adapter within a fine tuned Stable Diffusion UNet. To further reinforce identity fidelity, we propose a pseudo discriminator loss based on ArcFace cosine similarity with exponential timestep weighting. Experiments on held out and unseen faces show that Diff-ID does not exceed InstantID in raw ArcFace Face Similarity, but achieves substantially lower FID and the strongest FIQ based identity--realism trade off among the evaluated methods. We also present a unified DDIM based morphing pipeline that enables qualitative facial interpolation without per identity fine tuning. We further argue that identity preservation and photorealism should be evaluated jointly rather than in isolation, as high identity similarity alone does not guarantee realistic outputs. To make this trade off explicit, we report Face Image Quality (FIQ) as a complementary ratio based score that combines identity similarity and perceptual realism while keeping FS and FID as the primary metrics.

T. Rizwan, Sara Atito, Muhammad Awais et al. · 0 citations
Conference Open access Aug 2026

XSA-Mad: Cross-Modal Semantic Alignment for Morphing Attack Detection

Morphing attacks pose a serious threat to face recognition systems. However, existing image-based morphing attack detection (MAD) methods often generalize poorly to unseen generation techniques because they rely solely on visual cues. We propose XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces. Morphing concepts are decomposed into four interpretable attributes, including identity, facial geometry, texture, and consistency, and are encoded as structured and attribute-aware textual representations. The image encoder is progressively aligned with this discriminative textual space, resulting in a unified semantic representation that captures generation-invariant and concept-level discrepancies between bona-fide and morph images. Experiments on MAD22 and MorDIFF, following training on SMDD, demonstrate strong generalization across diverse morphing principles. In particular, XSA-MAD achieves an equal error rate of 2.92% on GAN-based morphs and consistently outperforms existing methods under high-fidelity generative attacks.

Jie Jin, Mahiro Tokumasu, Yushi Makino et al. · 0 citations
Open access Jul 2026

A Django-Enabled Hybrid Framework for Intelligent Face Morph Synthesis and Authentication Resilience

Facial recognition systems are widely used for identity verification but are vulnerable to face morphing attacks, where multiple facial images are blended to form a deceptive identity that can fool recognition models. This project develops a deep learning-based approach to detect such attacks and strengthen biometric authentication systems. It lies in the domain of Artificial Intelligence and Machine Learning, focusing on Computer Vision techniques to differentiate real and morphed facial images for accurate and secure verification. The project involves creating realistic morphed face datasets and building an efficient detection model applicable to border control, ID verification, and digital authentication. Current systems fail against high-quality morphs produced using advanced tools, showing reduced accuracy under variations in lighting, age, and facial accessories. To overcome this, the proposed model combines deep learning-based feature extraction with machine learning classifiers. Morph-2 and Morph-3 datasets are generated using professional morphing tools, and image enhancement with feature fusion is applied to improve accuracy and robustness.

Mekala Pooja, N. N. Kumar · 0 citations
Open access 2026

Decoherence-Aware Quantum State Evolution With Identity-Liveness Consistency for Robust Face Anti-Spoofing

Face anti-spoofing is essential for protecting biometric systems from presentation attacks such as print, replay, cut-photo, and deepfake manipulations. This work proposes a Decoherence-Aware Quantum State and Liveness Detection Framework for robust live-spoof discrimination. Facial video frames are first preprocessed and then passed through multi-scale convolutional networks to capture fine-grained spoof traces, such as moiré patterns, reflection noise, and display artifacts. The extracted features are then transformed into latent temporal states using superposition modeling and reversible temporal evolution. A stable evolution-adaptive transition mechanism is introduced to detect temporal disturbances induced by spoofing. Further, identity-liveness dependency stability modeling improves robustness against identity-preserving attacks. Quantum uncertainty-guided anomaly scoring is used for final spoof discrimination, followed by a lightweight MLP classifier. Experimental results on the CelebA-Spoof, OULU-NPU, and CASIA-FASD benchmark datasets demonstrate that the proposed framework achieves consistent performance across these evaluated datasets. The reported robustness and generalization are supported within the scope of these benchmark evaluations. On CelebA-Spoof, the model achieved APCER of 1.78%, BPCER of 2.29%, and ACER of 2.04%. On OULU-NPU, APCER, BPCER, and ACER were 1.38%, 1.42%, and 1.40%, respectively. On CASIA-FASD, the framework achieved 96.99% multiclass spoof classification accuracy, confirming better robustness and generalization.

D. Ameenulhakeem, O. N. Uçan · 0 citations
Preprint Jul 2026

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

The proposed DiffAttack framework significantly outperforms existing adversarial techniques, achieving a high average attack success rate of 84.86% across multiple face recognition models (e.g., FaceNet).

Omid Ahmadieh, Nima Karimian · 0 citations