Conceptual

Audio-Only Lightweight CNN for Touch-Gesture and Emotion Recognition in HRI

A privacy-preserving approach to human-robot interaction that recognizes touch gestures and emotional state from the sound a touch makes rather than from tactile skin or facial video. It introduces MTRCNN, a lightweight multi-temporal-resolution convolutional network using Mel-filterbank features to classify arousal and valence and distinct gestures, matching much larger pretrained audio networks at a tiny fraction of the parameters, size, and computation.