Conceptual

SeFAR: Semi-Supervised Fine-Grained Action Recognition

A semi-supervised framework for fine-grained action recognition (distinguishing subtle action variants within short temporal spans). It represents video with dual-level temporal elements, applies moderate temporal perturbation as a strong augmentation inside a Teacher-Student paradigm, and adds Adaptive Regulation to stabilize learning against the high uncertainty of teacher predictions on fine-grained classes. Reaches state-of-the-art on FineGym and FineDiving and improves coarse-grained UCF101/HMDB51, with features that also boost multimodal foundation models' fine-grained understanding.