Conceptual

Image-to-Video Surrogate Attacks for Transferable Adversarial Examples on Video Multimodal LLMs

A black-box attack that crafts adversarial videos able to fool unseen video multimodal language models. It uses an image-based multimodal model as a surrogate, perturbing frames by projected gradient descent to disrupt the latent video representation while fusing textual and spatiotemporal cues, and propagates perturbations across frames to survive unknown frame-sampling schemes.