K
KITT
Text
An efficient Transformer-based vision architecture (TDFN) that pairs a low-resolution global channel with a high-resolution channel steered by a reinforcement-learned fixation-point generator, so costly high-resolution processing is applied only to task-relevant regions of interest, cutting computation while preserving classification accuracy.