Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection
FakeSTormer detects manipulated (deepfake) videos in a way that generalizes to unseen generation methods by replacing a single real-versus-fake classifier with a multi-task spatio-temporal transformer whose auxiliary branches localize the spatial regions and temporal frames most vulnerable to artifacts. A video-level self-blending synthesis strategy manufactures pseudo-fake clips with subtle, automatically annotated artifacts to supervise those branches, improving cross-dataset detection robustness.
Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection Dat
Presents FakeSTormer, a method for detecting deepfake videos that generalizes to manipulation techniques unseen during training. Rather than a single real-versus-fake binary classifier, it uses a mul…