Conceptual

Attack-as-Defense Backdoor Injection Against Model Extraction in Deep Learning

A defense paradigm for protecting black-box classifiers from model extraction: instead of only blocking or watermarking, the victim exposes a honeypot output layer whose poisoned probability vectors implant a backdoor into any substitute model an attacker trains, enabling ownership verification and a trigger-driven reverse attack.