Conceptual

Modular Detection-Augmented Bandit Procedures for Piecewise-Stationary Multi-Armed Bandits

A modular framework for tackling multi-armed bandits whose arm-reward distributions stay fixed between abrupt change-points. It decouples the design into any stationary bandit algorithm plus any sequential change detector satisfying stated properties, so a single regret analysis yields order-optimal dynamic regret across many detector-plus-bandit combinations under sub-Gaussian rewards. Forced exploration guarantees that shifts even in rarely-pulled arms are detected, and improved lower bounds characterize the achievable dynamic regret.