XSA-MAD: Cross-modal Semantic Alignment for Morphing Attack Detection

Jie Jin, Mahiro Tokumasu, Yu Makino, Masakatsu Nishigaki, Tetsushi Ohki
Proceedings of the IEEE International Conference on Image Processing (ICIP), 2026.
[ Paper ] [ Web ]

Abstract

Morphing attacks pose a serious threat to face recognition systems. Existing image-based morphing attack detection methods often generalize poorly to unseen generation techniques because they rely solely on visual cues. XSA-MAD is a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces. Morphing concepts are decomposed into four interpretable attributes: identity, facial geometry, texture, and consistency. The image encoder is progressively aligned with this discriminative textual space, resulting in a unified semantic representation that captures generation-invariant, concept-level discrepancies. Experiments on MAD22 and MorDIFF, following training on SMDD, demonstrate strong generalization across diverse morphing principles. XSA-MAD achieves an equal error rate of 2.92% on GAN-based morphs.

Updated: