K
KITT
Text
Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach
Video-based multimodal large language models (V-MLLMs) answer questions about videos but can be fooled by adversarial videos -- clips altered with small, optimized pixel perturbations that change the…