Conceptual

Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection

FakeSTormer detects manipulated (deepfake) videos in a way that generalizes to unseen generation methods by replacing a single real-versus-fake classifier with a multi-task spatio-temporal transformer whose auxiliary branches localize the spatial regions and temporal frames most vulnerable to artifacts. A video-level self-blending synthesis strategy manufactures pseudo-fake clips with subtle, automatically annotated artifacts to supervise those branches, improving cross-dataset detection robustness.