Audio-Only Lightweight CNN for Touch-Gesture and Emotion Recognition in HRI
A privacy-preserving approach to human-robot interaction that recognizes touch gestures and emotional state from the sound a touch makes rather than from tactile skin or facial video. It introduces MTRCNN, a lightweight multi-temporal-resolution convolutional network using Mel-filterbank features to classify arousal and valence and distinct gestures, matching much larger pretrained audio networks at a tiny fraction of the parameters, size, and computation.
2501.00038
This ICASSP 2025 paper (cs.HC/cs.SD) proposes recognizing touch gestures and emotions during human-robot interaction from the sounds produced by touching a robot, avoiding both the need for full-body…