Back to papers
arxiv6.0 / 10

FIDA: Feature Instability-Driven Attack on Self-Supervised Facial Representation

Zhiyang Chen, Changchun Yin, Huiqin Yang, Liming Fang

Abstract

Self-supervised learning (SSL) models are vulnerable to backdoor attacks. However, the systemic risks they pose in face representation have received little attention. The entanglement of identity features in self-supervised face learning presents unique challenges for attack stealthiness. To address this gap, we propose FIDA (Feature Instability-Driven Attack), a novel backdoor attack framework. FIDA uses subtle semantic triggers for injection, but its key innovation is a novel objective called Feature Instability Loss. It trains the encoder to increase the sensitivity of triggered features along perturbation directions sampled during attack optimization . By preventing the backdoor from exhibiting the rigid feature patterns typical of previous attacks, FIDA effectively evades the evaluated perturbation-based defenses. Experiments show that FIDA achieves a high attack success rate and generally preserves benign utility across the evaluated settings , posing a significant threat to real-world multimedia applications relying on facial analysis.

Research area

ai securityrobustnesssecurity
Published
27 Aug 2026
Source
arxiv
Org
Nanjing University of Aeronautics and Astronautics
View paper
Sign in to read and join the discussion.