Conceptual
Login

Single-Agent Backdoor Leverage Attacks on Cooperative Multi-Agent RL

How a backdoor implanted in just one agent of a cooperative multi-agent deep reinforcement learning team can pry the entire system into failure. Covers spatiotemporal behavior patterns as stealthy triggers spread across a sequence of observations instead of instant visual patterns, decoupling the trigger from a controllably delayed attack period, and reward hacking with a unilateral influence filter that amplifies the backdoored agent's influence on teammates while suppressing the reverse influence. Includes target-failure-state guidance mined from historical trajectories, attack results against VDN, QMIX, and MAPPO teams, and resistance to representative backdoor defenses.